注册并分享邀请链接,可获得视频播放与邀请奖励。

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023–Deep Time «C’est la guerre.» ®1
加入 September 2010
3.3K 正在关注    77.1K 粉丝
One difference between V4 and V4.1 papers is the details on sparse attention training. They say much more. We know concretely that they pretrain with 64K for 34T tokens, and then do 11T at 1M. No dense warm-up. No instabilities throughout.
显示更多