Register and share your invite link to earn from video plays and referrals.

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
Joined January 2026
325 Following    3.7K Followers
It was one of the coolest, most thrilling tech adventures I’ve ever done. I got a wild buzz.
Qwen3.8 Flash Next just served a 262k token prompt with an image on my 4x3090. Decode at that depth: 67 tok/s. At 260K its 66.3, so the curve is basically FLAT. Almost every number you see for this model is 20-21 tok/s on a single 24GB card with experts offloaded to RAM on llamacpp. I went the other way: vLLM, all experts on GPU, FP8 KV. Full recipe below in my github, one command, raw runs included. 🧵A deep-dive technical thread. Let's go!
Show more