Top Models For your Hardware 2026
-- 8GB --
- lfm-2.6B: very solid, trained on tons of data, obviously limited by good place to start
Use this for snippet tool calls, or to play around and build infra
-- 16GB --
I am a big fan of Ornith models, they do well at tool calling and they certainly push models farther
Gemma is a more well rounded model for chat, and vision IMO
-- 24GB up to 96GB --
Qwen3.8-27B is IMO the first model here you can code with, it's really strong. Use the exl3 versions.
-- 96GB up to 196GB --
Qwen3.8-Flash-Next is where it's at. Phenomenal model, very fast even on slower hardware, kv cache is smaller and part of the model is basic enough to be offloaded to ram
-- 196GB up to 384GB --
GLM-5.3-Flash is frontier at home, built to run on 10,000$ of hardware. It's really a gift. It is natively multimodal, image/video/ audio?
It hold up over 1 million tokens in context, and is hybrid attention, it holds on higher concurrency. Very good for 3D, coding, hacking, art.
-- 384GB up to 512GB --
Sacrifice speed for intelligence, GLM-5.3 is tied for #
1# Open Weight model, and is the best for pure coding and systems work.
I told you August was going to be awesome, now we go to slower months.
-------------
It's not Local AI we have to worry about anymore.