We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization.
@deepseek_ai stan #1#, 2023โDeep Time
ยซCโest la guerre.ยป ยฎ1
One difference between V4 and V4.1 papers is the details on sparse attention training. They say much more.
We know concretely that they pretrain with 64K for 34T tokens, and then do 11T at 1M.
No dense warm-up. No instabilities throughout.