Register and share your invite link to earn from video plays and referrals.

Search results for Prune
Prune community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Prune
Prune Juice Was Served On A Packed Flight With Just Two Lavatories. What Could Go Wrong?
benchmarks of a 50% pruned Qwen3.6-35b-a3b and expert-specific quantization technique (made by me) 7.3gb model preforming => 51gb model, exiting to see where I can bring this technique to. I have some more things lined up too. I need a DGX spark😭
Show more
Kijai's 😃 Minimax_h3_ref2va_pruned_w6a8_g32 its experimental, highly quantized work-in-progress (WIP) base model for MiniMax H3 video generation. It is a massive 16.8 GB file designed to drastically reduce the VRAM required to run this model locally. 👇
Show more
GM. today's agenda: prune, measure, repeat. 🪴
Kijai's 😃 Minimax H3 4step_lora_flashgen_v1.0 768p_fl2va_pruned_avg_rank_13_bf16 This Lora is a Turbo/Flash acceleration adapter. using only 4 steps, delivering a roughly 5x speedup in generation times 👇
Show more
Updated papercuts and DCE to 0.3.0 Papercuts: you can now prune the papercuts list after you resolve DCE: the agent asks you at the end of a session if any of the tools that were promoted should be alwaysActive. pi install npm:pi-papercuts pi install npm:pi-deferred-context-engine
Show more
Qwen3.8-2.4T-A95B, compressed two ways at once: 25% of the experts pruned with REAP, and the rest quantized to NVFP4. Even with a quarter of the experts gone and 4-bit weights, GPQA Diamond holds at 91.5 vs 92.6 for the full-precision base. ~99% recovery. Serve on @vllm_project:
Show more
H3 🐰💋👠 N$FW LoRA -PinkFluffyBunny lora that truly brings the style and pose -Maximum results achieved at 0.5 str on pruned int8 model. -Use the h3_fl2va_pruned_int8_convrot.safetensors model. Alpha quality so temper expectations 👇😉
Show more
MiniMax H3 💋💄👠Naughty Times LoRA Tip- Run it with a lora loader at value 0.5 using the model (minimax_h3_fl2va_pruned_int8_convrot.safetensors) 👇
OpenAI’s Navier–Stokes result reportedly used ~10,000 autonomous agents, but only 2.7M messages and 130B tokens for the final proof search. That’s not “lots of agents = lots of chatter.” It’s tight coordination: orchestration logic aggressively prunes, routes, and reuses partial work.
Show more