๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
๊ฐ€์ž… June 2025
10 ํŒ”๋กœ์ž‰ ์ค‘    11.5K ํŒฌ
๐Ÿšจ BREAKING: these engineers figured out how to serve Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream! this crushes B200 by 3.8x in aggregate throughput/node and 1.3x in single stream decode + beats B300 on performance per dollar (48 vs 33 tok/s/$) See how in the thread.
๋” ๋ณด๊ธฐ