註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
"Matryoshka Language Model Suites" Instead of training every model size separately, this paper nests 500M, 1.5B, and 3B models inside one architecture and trains them together. The smaller models are standalone checkpoints, get near-free distillation from the largest model, and share weights + KV cache for speculative decoding. And you still get the same performance, with 36% less training compute, and 14-26% faster speculative decoding. So a model family can become one jointly trained system instead of several independent models.
顯示更多
0
4
396
53
轉發到社區