Register and share your invite link to earn from video plays and referrals.

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
Joined January 2026
325 Following    3.7K Followers
I might fix a 180B model with 8.4 MiB. Thats 72 tiny FP8 weights in my Qwen3.8-Flash-Next Franken-quant They live in the part that writes model state token by token. If FP8 rounds badly there, the error doesnt disappear. It gets fed back into the next token, then the next one, and a long conversation can slowly go weird. So Im restoring only in_proj_a and in_proj_b in Gated Delta Net to their original BF16. 8.4 MiB per 3090. 262K context still fits. AWQ INT4 experts, FP8 KV, Vision and CUDA graphs stay exactly where they are. Maybe this does nothing. Maybe 72 tiny tensors are the part that was quietly making the model dumber. Same tests on both builds. Vision, long context, multi-turn, agent tasks. Then we find out.
Show more
YES, FUCKING YES! I DID IT !!! In my last post, I heavily trashed Qwen3.8 Flash Next, claiming it couldn't do anything at all. I managed to get awesome speed out of it, but in nearly 3 hours it couldn't even generate a simple Pagoda Garden. I had no clue what was going on. Turns out the issue was with my custom quant, which I built myself to fit 256k context + vision into 4x3090 with good throughput. @_cpatonn released a properly calibrated AWQ INT4 of this model, but it was larger than mine and didn't fit into my 96GB as-is. CUDA OOM - sounds familiar to everyone, right? I decided to tackle it and make something that actually works. A full day of work, interrupted only by 5 hours of sleep. To fit 262K context, I had to un-fuse the fused layers and quantize all non-sensitive projections to FP8. vLLM allows only 1 quantization scheme per fused layer, which turned into writing a custom loader and 13 launch attempts 😭 Attempt 12 finally loaded, and I was so hyped. But the model's output was pure gibberish, total mush. Turns out TP4 was slicing the fused QKV matrix in the wrong places, for fuck's sake. Attempt 13 fixed the slicing. 262K fit. The model kept spitting garbage on any complex task. Short answers were fine, a 40K needle test was fine, but ask for a three.js scene - and you get 16,384 tokens of complete shit. I was devastated. The root cause of all this was just mind-bogglingly fucked up. 48 tiny tensors, shared_expert_gate.weight, ~5 KB each. My build script quantized them to FP8, but the runtime read them as BF16 without applying the scale 🤡 The gate expects values around 0.077, but was receiving 448 🤡🤡 Eat up, my dear friend. And this gate determines how much the shared expert contributes to every single token. I restored all 48 from the original checkpoint - and the pagoda scene, which I had spent hours grinding on and ran dozens of times - just passed twice in a row. Wanna know how mindblown I was? Exactly. Then the model swallowed 258,008 real prompt tokens with an image and nailed the needle test. Decode at 63-68 tok/s at 220 W per GPU - identical to what the broken build pushed. Video is the proof. One prompt, two sets of weights: reference BF16 vs. my Franken-quant running on 4 gaming GPUs from 2020. One-shot, zero fixes. Mine finished first, and the scene came out tighter. It's just 1 prompt, so I won't call it a benchmark. 24 hours of my life for 240 KB of broken tensors 🤡🤡🤡
Show more