注册并分享邀请链接,可获得视频播放与邀请奖励。

Niels Rogge
@NielsRogge
ML Engineer @huggingface. Building @KU_Leuven grad. General interest in machine & deep learning. Making AI more accessible for everyone!
加入 April 2010
744 正在关注    22.5K 粉丝
For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention (SWA): A local attention pattern that restricts each token to attending only within a fixed-size neighborhood instead of the full sequence. This reduces attention and KV-cache memory for long-context models, while periodic global-attention layers can preserve broader context. Find it here:
显示更多
0
3
206
29
转发到社区