Register and share your invite link to earn from video plays and referrals.

Search results for mtp
mtp community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including mtp
Telegram MTProto proxies are now live in the TorGuard V2Ray panel. Manage MTProto proxies alongside your V2Ray configs for another tool to bypass web censorship:
llama.cpp with MTP support makes local models fast enough to use as daily drivers 🚀 Qwen3.6-27B dense generation below on A10G: From 25 tok/st to 45 tok/s (+78%)!
We released experimental MTP Qwen3.6 Unsloth GGUFs! Qwen3.6 27B MTP now runs at 140 tokens/s. Qwen3.6 35B-A3B MTP gets 220 tokens/s generation on a single GPU. Qwen3.6 27B and 35B-A3B have >1.4x speed-up over the original GGUFs without any change in accuracy. Guide + GGUFs + Benchmarks: In terms of average speedup, we see a 1.4x for dense models at draft tokens = 2 and for the MoE around 1.15 to 1.2x. We do not recommend more than 2 draft tokens because the acceptance rate drops precipitously from 83% to 50% with 4 draft tokens, and the forward passes for MTP become less beneficial. Use `--spec-type mtp --spec-draft-n-max 2` Thanks to Aman for
Show more
0
60
764
113
Forward to community
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide:
Show more
0
62
2.2K
259
Forward to community
you gotta guard kat from everywhere 😵‍💫
Benchmarking @NVIDIAAI's Nemotron Puzzle 75B locally on the GX10. NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations. 🏃‍♀️22.75 tok/s in a single session into 88.85 cumulative at 7 sessions. 🧍Baseline without MTP: ~16.3. Scripts + setup:
Show more