가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

alphaXiv
@askalphaxiv
High fidelity research
가입 November 2023
101 팔로잉 중    56.5K 팬
"Matryoshka Language Model Suites" Instead of training every model size separately, this paper nests 500M, 1.5B, and 3B models inside one architecture and trains them together. The smaller models are standalone checkpoints, get near-free distillation from the largest model, and share weights + KV cache for speculative decoding. And you still get the same performance, with 36% less training compute, and 14-26% faster speculative decoding. So a model family can become one jointly trained system instead of several independent models.
더 보기