Register and share your invite link to earn from video plays and referrals.

Sami Kassab
@Old_Samster
Lore stacking @UnsupervisedCap | pushin' τ
Joined February 2013
1.1K Following    18.3K Followers
Two things: - Quasar is a scam - stay away. They’ve been lying since day 1. I was one of the first bittensor investors they pitched and after some diligence, we found out they were lying about their past experience. I told everyone who’d asked me for my opinion afterwards to stay away. - Do your own damn research and stop trusting people Here’s a tweet our team put together on Friday after the novelty search that we were saving for Monday: —————————- There are a lot of big claims Quasar is making that are hard to verify with the current public information available. Given the incredibly difficult work that goes into decentralized model training, we’re hoping the Quasar team can provide more clarity around the research and work they’re doing. Below are some of the areas that require further clarification. First, the 5M context window is a big claim and one of the core differentiators with their model, but none of the benchmarks released by the team evaluate the model on long context abilities. Quasar’s architecture substantially compresses context, meaning that, although the model can hypothetically process 5M tokens, the proof that it can stay coherent requires benchmarking on standard long context evals like RULER and/or NIAH. We’re awaiting the results from these benchmarks before we evaluate the subnet’s progress on this claim. Second, the training process itself is unclear. True decentralized training means using untrusted and non-collocated compute. Was the 120B model that was trained on 7T tokens trained in a centralized setting or in a decentralized setting using the subnet? If it was trained on the subnet, is there any data showing that the training process is fault tolerant and can handle interruptible nodes? Which techniques are the team using to compress the data (e.g., gradients, activations, or both) shared across miners? How many independent compute contributors participated or are participating in the training run? What are the hardware requirements for compute contributors?
Show more