๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

MiniMax (official)
@MiniMax_AI
Code: @MiniMaxAgent MiniMax Design: @Hailuo_AI API: Token Plan:
๊ฐ€์ž… January 2025
994 ํŒ”๋กœ์ž‰ ์ค‘    124.1K ํŒฌ
Three paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Hereโ€™s a visual guide. ๐Ÿ‘‡ Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? Weโ€™re looking forward to seeing that explored.
๋” ๋ณด๊ธฐ