Register and share your invite link to earn from video plays and referrals.

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
Joined January 2026
325 Following    3.7K Followers
Every demo in this video came out of a single prompt. No follow ups, nothing fixed by hand. I ran 168 of them through an agent harness, one shot each, and picked 20 for the cut. Median 7 minutes per demo. The neon tunnel burned 98 056 output tokens to produce 5 KB of code, it thinks a lot more than it writes. 68 of the 168 came out working in a browser. Most of the misses were my own harness capping a reply at 64 000 tokens, not the model. This is Hy4 preview, Tencent Hunyuan just open sourced it. 770B total, 49B active, 1M+ context, third flagship they ship in 6 months. The part i find more interesting than the size is how they built it. They co-designed it next to their own products instead of throwing the weights over the wall, with people who actually do software, games, finance and security. In their internal blind test, 163 experts across 203 engineering tasks, it scored 2.99/4 against Kimi K3 at 2.94 and GLM 5.3 at 2.92. Their numbers, not mine. Price is the part that will annoy some people. $0.834/M in, $2.501/M out, $0.042/M cache hits. Hy4 preview is free on WorkBuddy for 2 weeks right now if you want to poke at it yourself: @TencentHunyuan @TencentAI_News @WorkBuddy_AI
Show more