註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Fireworks
@FireworksAI_HQ
The frontier platform for training and inference on open-weights models at scale.
加入 September 2022
281 正在關注    29.9K 粉絲
Long-context sparse attention has a catch: data-dependent block selection wrecks memory access kills speed. Our @MiniMax_AI M3 kernel on Blackwell answers it. KV-stationary, each block read once, ~980 TFLOP/s on a B200. See the breakdown here →
顯示更多