注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
75 正在关注    51.4K 粉丝
What if you have transformers learn to think in both latent states & tokens? This Microsoft paper adds a self-supervised next-latent prediction objective to standard next-token training, where a lightweight dynamics model learns to predict the transformer’s next hidden state from the current hidden state and next token. This pushes transformer representations toward compact belief states with coherent transition dynamics, improving world modeling, reasoning, planning, and enabling variable-length self-speculative decoding with up to 3.3x inference speedup.
显示更多
0
6
186
27
转发到社区