๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alexey Fateev
@superalesha
โšกI benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. โค๏ธ - 2xDGX Spark ๐Ÿš€96GB VRAM | Local AI
๊ฐ€์ž… January 2026
325 ํŒ”๋กœ์ž‰ ์ค‘    3.7K ํŒฌ
Qwen3.8 Flash Next just served a 262k token prompt with an image on my 4x3090. Decode at that depth: 67 tok/s. At 260K its 66.3, so the curve is basically FLAT. Almost every number you see for this model is 20-21 tok/s on a single 24GB card with experts offloaded to RAM on llamacpp. I went the other way: vLLM, all experts on GPU, FP8 KV. Full recipe below in my github, one command, raw runs included. ๐ŸงตA deep-dive technical thread. Let's go!
๋” ๋ณด๊ธฐ