登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research
参加 November 2023
101 フォロー中    56.5K ファン
"Matryoshka Language Model Suites" Instead of training every model size separately, this paper nests 500M, 1.5B, and 3B models inside one architecture and trains them together. The smaller models are standalone checkpoints, get near-free distillation from the largest model, and share weights + KV cache for speculative decoding. And you still get the same performance, with 36% less training compute, and 14-26% faster speculative decoding. So a model family can become one jointly trained system instead of several independent models.
もっと見る