註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023–Deep Time «C’est la guerre.» ®1
加入 September 2010
3.3K 正在關注    77K 粉絲
One difference between V4 and V4.1 papers is the details on sparse attention training. They say much more. We know concretely that they pretrain with 64K for 34T tokens, and then do 11T at 1M. No dense warm-up. No instabilities throughout.
顯示更多