注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research. Dm for promo
加入 November 2023
72 正在关注    50.7K 粉丝
OCR but with a working memory...?! This paper, Unlimited OCR, replaces decoder full self-attention with Reference Sliding Window Attention, giving the model a working memory where each token attends to the fixed visual and prompt references plus only the most recent output tokens, making decode KV cache constant instead of linear in generation length. Combined with DeepEncoder’s 16x visual compression, this working-memory-style attention enables one-shot multi-page OCR with stable memory and latency, while improving OmniDocBench performance over DeepSeek OCR.
显示更多
0
8
482
70
转发到社区