註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Niels Rogge
@NielsRogge
ML Engineer @huggingface. Building @KU_Leuven grad. General interest in machine & deep learning. Making AI more accessible for everyone!
加入 April 2010
744 正在關注    22.5K 粉絲
For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention (SWA): A local attention pattern that restricts each token to attending only within a fixed-size neighborhood instead of the full sequence. This reduces attention and KV-cache memory for long-context models, while periodic global-attention layers can preserve broader context. Find it here:
顯示更多
0
3
206
29
轉發到社區