가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Cheng
@zcbenz
maintainer of MLX @apple. creator of @electronjs. check for the open source things I built.
가입 June 2007
107 팔로잉 중    7.6K 팬
For linear attention a typical implementation only caches the last hidden states, i.e. the context of last token. While for speculative decoding like MTP we need to rollback at least n_draft tokens' context, which does not work with linear attention's cache. An intuitive solution is to remember at least n_draft hidden states in linear attention's cache, I experimented with a cache implementation (which is a bit frustrating to write) and I think the idea works well. The really confusing thing is, no one seems to take this approach, i.e. making a general cache for linear attention that stores hidden states in temporal manner with fixed size, instead people wire MTP inside GDN with a checkpoint of the hidden states during the draft window. Surely it works, but isn't it the ugliest possible implementation and you would have to pollute code of every model that uses MTP? I can see an answer is to avoid increased RAM usage, but is is really small keeping only n_draft context, and the intrusive changes to model implementation kill the elegant abstractions.
더 보기