註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Anjney Midha
@AnjneyMidha
founder @amppublic • visiting scientist @stanford • teaching @cs153systems • alignment was always the moat.
加入 May 2011
3K 正在關注    48.6K 粉絲
very cool a 2-3x speed up in training by essentially letting the model learn more flexibly in its early stages than rigid regimes sort of akin to how homeschooling is much better for some kids than factory education
顯示更多
Today we release Token Superposition Training (TST), a modification to the standard LLM pretraining loop that produces a 2-3× wall-clock speedup at matched FLOPs without changing the model architecture, optimizer, tokenizer, or training data. During the first third of training, the model reads and predicts contiguous bags of tokens, averaging their embeddings on the input side and predicting the next bag with a modified cross-entropy on the output side. For the remainder of the run, it trains normally on next-token prediction. The inference-time model is identical to one produced by conventional pretraining. Validated at 270M, 600M, and 3B dense scales, and at 10B-A1B MoE. The work on TST was led by @bloc97_, @gigant_theo, and @theemozilla.
顯示更多
0
6
92
10
轉發到社區