註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

MiniMax (official)
@MiniMax_AI
Code: @MiniMaxAgent MiniMax Design: @Hailuo_AI API: Token Plan:
加入 January 2025
994 正在關注    124.1K 粉絲
Three paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.
顯示更多
0
20
329
35
轉發到社區