Register and share your invite link to earn from video plays and referrals.

Search results for speeditup
speeditup community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including speeditup
Since when do you slow down 20 miles when it rains #speeditup#
Unlocking Lossless Speedups in LLMs via Discrete Diffusion paper:
Finally, a speedup method for Minimax H3 reference to video! ComfyUI-Ref2VA-VSA adds video sparse attention to Minimax, generating 5-second clips in ~72s on an RTX 4090. It keeps reference tokens dense while pruning video attention—2.24× faster than VDN-H3 at ~13.5GB VRAM.
Show more
⚡ New RTX local-agent optimizations include a vLLM speedup on Blackwell. @NVIDIARTXSpark reports: 🛠️ 1.2x vLLM performance on RTX PRO 6000 Blackwell ⚡ Up to 1.4x on a two-system DGX Spark cluster Great to see the local serving path getting faster across RTX and DGX Spark.
Show more
Lossless quantization has usually meant giving up inference speedup. This paper changes that. SLQ (Statistically-Lossless Quantization) reaches task-lossless compression at 3.3 bits per parameter, and distribution-lossless at 5-6 bpp where the output distribution is practically indistinguishable from the original. 1.7 to 3.6x throughput over BF16 in @vllm_project. Beats FP8 while staying lossless. From Michael Helcig, @_EldarKurtic, and @DAlistarh.
Show more
🚀 Introducing FlashQLA: high-performance linear attention kernels built on TileLang. ⚡ 2–3× forward speedup. 2× backward speedup. 💻 Purpose-built for agentic AI on your personal devices. 💡Key insights: 1. Gate-driven automatic intra-card CP. 2. Hardware-friendly algebraic reformulation. 3. TileLang fused warp-specialized kernels. FlashQLA boosts SM utilization via automatic intra-device CP. The gains are especially pronounced for TP setups, small models, and long-context workloads. Instead of fusing the entire GDN flow into a single kernel, we split it into two kernels optimized for CP and backward efficiency. At large batch sizes this incurs extra memory I/O overhead vs. a fully fused approach, but it delivers better real-world performance on edge devices and long-context workloads. The backward pass was the hardest part: we built a 16-stage warp-specialized pipeline under extremely tight on-chip memory constraints, ultimately achieving 2×+ kernel-level speedups. We hope this is useful to the community!🫶🫶 Learn more: 📖 Blog: 💻 Code:
Show more
0
33
1.3K
149
Forward to community
LLM Compressor v0.14.0 is out, and GPTQ just got its biggest speedup since launch. A new Triton kernel makes quantization ~15x faster end to end. Batching layers that share a shape pushes that to ~30x on some MoE workloads. Even the old eager path is 1.5-2x faster. Also new: expanded MSE/iMatrix observers that beat GPTQ for NVFP4 on internal benchmarks, REAP pruning with distributed DDP, and support for GLM 5.3 and Qwen3.8. Full release notes:
Show more
🚀 MiniMax H3 Super Acceleration in Sol-Engine🤩 We pushed H3 far beyond our previous 3–4× optimization regime — reaching 22.2× speedup for 5s video and 27.7× for 10s video vs. the published SGLang baseline. On a single NVIDIA GB200: • 5s 768p: 152.3s → 6.85s (22.2×) • 10s 768p: 414.1s → 14.93s (27.7×) But the more interesting part may be what this means economically. Using MiniMax’s published H3 API price as a reference, we translate inference speed directly into production economics. Under an ideal fully utilized GB200 scenario, Sol-Super can serve about 525 five-second videos/hour, corresponding to roughly $210/hour of output value at the reference API price. Assuming $5.50/GPU-hour, that implies a 97%+ GPU-only gross margin in the idealized model. At full utilization, one GB200 could produce: • 12.6K 5s videos/day • 378K videos/month • equivalent to 525 hours of finished video per month For us, this is the bigger point of inference optimization: a 20×+ speedup does not just reduce latency — it can fundamentally change the unit economics, serving capacity, and viable business models of video generation. 🔗
Show more
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results:
Show more
0
1K
39K
5.8K
Forward to community
Kijai's 😃 Minimax H3 4step_lora_flashgen_v1.0 768p_fl2va_pruned_avg_rank_13_bf16 This Lora is a Turbo/Flash acceleration adapter. using only 4 steps, delivering a roughly 5x speedup in generation times 👇
Show more