注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在关注    56.9K 粉丝
"Matryoshka Language Model Suites" Instead of training every model size separately, this paper nests 500M, 1.5B, and 3B models inside one architecture and trains them together. The smaller models are standalone checkpoints, get near-free distillation from the largest model, and share weights + KV cache for speculative decoding. And you still get the same performance, with 36% less training compute, and 14-26% faster speculative decoding. So a model family can become one jointly trained system instead of several independent models.
显示更多
0
4
396
53
转发到社区