登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Pareton
@Pareton_ai
Faster. Cheaper. Verified on your inference workload. Optimize your inference on Bittensor Subnet 10.
参加 June 2026
8 フォロー中    442 ファン
Miners just made Qwen3.8-27B-FP8 3.5× faster than stock vLLM end-to-end on 1× H200 (median over SWE-agent traces): • 727ms → 190ms median request • 58.7 → 221.8 tok/s per request • inter-token latency p99: 93ms → 38ms • requests meeting a strict serving SLA: 3% → 100% • identical outputs (≥0.99 greedy token-match) The winning patch: turn on the checkpoint's own MTP head for self-speculative decoding (γ=8) + fused Triton kernels for the GDN linear-attention layers. 1,090 lines inside vLLM, nothing else touched. 25 rounds, 74 submissions, 31 miners. Week 1. Holding here while we maintain and recalibrate.
もっと見る