Local AI people: what hardware stack actually replaces a $200/mo frontier coding subscription for a power user?
I keep seeing setups with 1, 2, 3, 4 DGX Sparks, and the OSS models are getting really good. But most benchmarks/use cases still seem focused on running one chatbot/agent at a time.
That’s not my workload.
I’m running 2-8, sometimes 10, Codex instances in parallel. I need frontier-ish coding quality AND enough throughput for every session to stay fluid.
Being able to fit a huge model is useless to me if it means ~20 tok/s and one useful session at a time.
So what’s the actual hardware stack for this?
Can stacked Sparks realistically handle that kind of concurrency, or does this basically require a proper multi-GPU setup?
Feels like most “local AI” discussions optimize for what you can run, while I care a lot more about how many instances I can run well at the same time.