Register and share your invite link to earn from video plays and referrals.

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023–Deep Time «C’est la guerre.» ®1
Joined September 2010
3.3K Following    76.9K Followers
One difference between V4 and V4.1 papers is the details on sparse attention training. They say much more. We know concretely that they pretrain with 64K for 34T tokens, and then do 11T at 1M. No dense warm-up. No instabilities throughout.
Show more