註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
加入 June 2025
10 正在關注    11.5K 粉絲
🚨 BREAKING: these engineers figured out how to serve Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream! this crushes B200 by 3.8x in aggregate throughput/node and 1.3x in single stream decode + beats B300 on performance per dollar (48 vs 33 tok/s/$) See how in the thread.
顯示更多
0
24
895
85
轉發到社區