Register and share your invite link to earn from video plays and referrals.

0xSero
@0xSero
permanent abundance
Joined December 2020
1.7K Following    71.8K Followers
Top Models For your Hardware 2026 -- 8GB -- - lfm-2.6B: very solid, trained on tons of data, obviously limited by good place to start Use this for snippet tool calls, or to play around and build infra -- 16GB -- I am a big fan of Ornith models, they do well at tool calling and they certainly push models farther Gemma is a more well rounded model for chat, and vision IMO -- 24GB up to 96GB -- Qwen3.8-27B is IMO the first model here you can code with, it's really strong. Use the exl3 versions. -- 96GB up to 196GB -- Qwen3.8-Flash-Next is where it's at. Phenomenal model, very fast even on slower hardware, kv cache is smaller and part of the model is basic enough to be offloaded to ram -- 196GB up to 384GB -- GLM-5.3-Flash is frontier at home, built to run on 10,000$ of hardware. It's really a gift. It is natively multimodal, image/video/ audio? It hold up over 1 million tokens in context, and is hybrid attention, it holds on higher concurrency. Very good for 3D, coding, hacking, art. -- 384GB up to 512GB -- Sacrifice speed for intelligence, GLM-5.3 is tied for #1# Open Weight model, and is the best for pure coding and systems work. I told you August was going to be awesome, now we go to slower months. ------------- It's not Local AI we have to worry about anymore.
Show more