가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Nunchux AI
@NunchuxAI
Building the frontier of multimodal inference for image, video, and world models.
가입 October 2025
3 팔로잉 중    566 팬
Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: Technical Report: Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.
더 보기