๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Teortaxesโ–ถ๏ธ (DeepSeek ๆŽจ็‰น๐Ÿ‹้“็ฒ‰ 2023 โ€“ โˆž)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023โ€“Deep Time ยซCโ€™est la guerre.ยป ยฎ1
๊ฐ€์ž… September 2010
3.3K ํŒ”๋กœ์ž‰ ์ค‘    76.9K ํŒฌ
One difference between V4 and V4.1 papers is the details on sparse attention training. They say much more. We know concretely that they pretrain with 64K for 34T tokens, and then do 11T at 1M. No dense warm-up. No instabilities throughout.
๋” ๋ณด๊ธฐ