Register and share your invite link to earn from video plays and referrals.

Pareton
@Pareton_ai
Faster. Cheaper. Verified on your inference workload. Optimize your inference on Bittensor Subnet 10.
8 Following    442 Followers
Pareton miners found it, vLLM merged it. An optimization from our Qwen campaign on #Bittensor# SN10 is now upstream in @vllm_project: ~4% more throughput at batch 4–8 for Qwen3.8 with MTP speculative decoding. Open competition → open-source wins. PR: 1/5
Show more
We just put Pareton's engine on an H200 and beat a B200 on Qwen3.8-27B-FP8. B200, stock vLLM: 106.2 tok/s H200, Pareton engine: 252.4 tok/s 2.4x on a cheaper GPU. Same prompt, same output length, batch 1. Third-party verified:
Show more
Miners just made Qwen3.8-27B-FP8 3.5× faster than stock vLLM end-to-end on 1× H200 (median over SWE-agent traces): • 727ms → 190ms median request • 58.7 → 221.8 tok/s per request • inter-token latency p99: 93ms → 38ms • requests meeting a strict serving SLA: 3% → 100% • identical outputs (≥0.99 greedy token-match) The winning patch: turn on the checkpoint's own MTP head for self-speculative decoding (γ=8) + fused Triton kernels for the GDN linear-attention layers. 1,090 lines inside vLLM, nothing else touched. 25 rounds, 74 submissions, 31 miners. Week 1. Holding here while we maintain and recalibrate.
Show more