注册并分享邀请链接,可获得视频播放与邀请奖励。

Dimitris Papailiopoulos
@DimitrisPapail
Researcher @Microsoft | Prof @UWMadison (on leave) | babas of Inez Lily.
加入 May 2012
1.5K 正在关注    30.1K 粉丝
Found something in my daily use of Claude Code that validates our Memento results: Claude Code flushes the KV cache after some idle period, and when I come back past that the model is noticeably harder to work with. Conjecture: post-flush, the model is no longer continuing its trajectory. It's shoved into a weird OOD regime where it has to simulate what has happened from the tokens and resume from a reconstruction. Which is much harder than just continuing!! We measured this effect in our paper. KV states (soft embeddings) carry information that text tokens don't, even when attention is masked. Bottom line: If you flush your cache you lose a lot of accuracy!
显示更多
0
42
827
70
转发到社区