Miners just made Qwen3.8-27B-FP8 3.5× faster than stock vLLM end-to-end on 1× H200 (median over SWE-agent traces):
• 727ms → 190ms median request
• 58.7 → 221.8 tok/s per request
• inter-token latency p99: 93ms → 38ms
• requests meeting a strict serving SLA: 3% → 100%
• identical outputs (≥0.99 greedy token-match)
The winning patch: turn on the checkpoint's own MTP head for self-speculative decoding (γ=8) + fused Triton kernels for the GDN linear-attention layers. 1,090 lines inside vLLM, nothing else touched.
25 rounds, 74 submissions, 31 miners. Week 1. Holding here while we maintain and recalibrate.