註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research. Dm for promo
加入 November 2023
72 正在關注    50.7K 粉絲
OCR but with a working memory...?! This paper, Unlimited OCR, replaces decoder full self-attention with Reference Sliding Window Attention, giving the model a working memory where each token attends to the fixed visual and prompt references plus only the most recent output tokens, making decode KV cache constant instead of linear in generation length. Combined with DeepEncoder’s 16x visual compression, this working-memory-style attention enables one-shot multi-page OCR with stable memory and latency, while improving OmniDocBench performance over DeepSeek OCR.
顯示更多
0
8
482
70
轉發到社區