註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Pareton
@Pareton_ai
Faster. Cheaper. Verified on your inference workload. Optimize your inference on Bittensor Subnet 10.
加入 June 2026
8 正在關注    442 粉絲
Miners just made Qwen3.8-27B-FP8 3.5× faster than stock vLLM end-to-end on 1× H200 (median over SWE-agent traces): • 727ms → 190ms median request • 58.7 → 221.8 tok/s per request • inter-token latency p99: 93ms → 38ms • requests meeting a strict serving SLA: 3% → 100% • identical outputs (≥0.99 greedy token-match) The winning patch: turn on the checkpoint's own MTP head for self-speculative decoding (γ=8) + fused Triton kernels for the GDN linear-attention layers. 1,090 lines inside vLLM, nothing else touched. 25 rounds, 74 submissions, 31 miners. Week 1. Holding here while we maintain and recalibrate.
顯示更多
0
4
59
12
轉發到社區