가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Turing Post
@TheTuringPost
On X we surface the AI research that matters and explain the ideas behind it. In the newsletter, we connect the dots between AI’s past, present, and future ⬇️
가입 June 2020
8.3K 팔로잉 중    88K 팬
This @GoogleDeepMind's paper is really worth your time It's on how to help Transformers not lose the right context on the way to the final answer. For this, the researchers introduce Recirculation: Normally, information passes through Transformer layers once. Recirculation changes that flow: → Some of what the model figures out in deeper layers is passed back to earlier layers and used when processing the next input. For example, once the model understands that “bank” means river bank in a fishing context, that interpretation can stay available when it later gets a question about an ATM. So here is how recirculation works: 1. The model processes the input normally. 2. Deeper layers build a more contextualized representation. 3. A small part of that activation is mixed back into a shallower layer. 4. The next input is processed with this updated state. The weights stay frozen, and you're changing how information flows through the model at inference time, not retraining it. This method really works in practice: - recirculation reduced contextualization errors by 60% - reduced perplexity by 23% - improved GSM8K accuracy by 21% - improved performance on several other tasks And now, the most interesting question: is this an alternative to Chain-of-Thoughts? Not really. CoT adds computation through generated reasoning tokens; recirculation helps the model keep track of what it has already understood internally. It’s also different from looped Transformers: they repeat the same layers, effectively adding depth, while recirculation feeds deeper representations back into earlier processing. So instead of asking the model to reason out loud, let it reuse more of what it has already figured out internally.
더 보기