가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Turing Post
@TheTuringPost
On X we surface the AI research that matters and explain the ideas behind it. In the newsletter, we connect the dots between AI’s past, present, and future ⬇️
가입 June 2020
8.3K 팔로잉 중    88K 팬
This is a new training strategy for model suits that you should try: → Train a whole family of models as one nested model with the Matryoshka framework. @Cornell researchers showed how they nested 3 models like Matryoshka dolls: • Each model builds on the smaller one. The 500M model sits inside the 1.5B model, which sits inside the 3B model. • A "junction" connects these models. It adapts the smaller model’s output to the larger one’s wider hidden representation, without adding new parameters. • They share parts of the same architecture and are trained in one run. One forward pass produces predictions from all model sizes. The smaller model does the first part of the computation, and the larger model can pick up from there instead of starting over. There are 2 very clear benefits of this method: 1. Smaller models learn from the largest model automatically. It’s built-in distillation. 2. Nesting helps speculative decoding. This benefit comes at inference time. A small model guesses the next tokens, and the larger one checks them. They share early layers, can reuse the KV cache, and also tend to agree more often. Compared with training all 3 models separately, Matryoshka suit achieves the same model quality while it: - uses 36% less training compute - makes speculative decoding 14–26% faster In any case, you can still pull out smaller models and use them independently.
더 보기