Here's all of the models currently in the "frontier" discussion for interactive, human in the loop coding sessions.
@ArtificialAnlys data.
Takeaways:
1. GLM-5.3-Flash is cheap but extremely slow (high output tokens per task)
2. Sol dominates the time per task frontier
3. Gemini 3.7 flash is very similar to Sol.
4. Opus is smart and fast, but costs far more.
5. Grok is good but a step down from Opus and Gemini.
6. OSS models require ~10-20x as many output tokens per task, which means they're extremely slow.
7. For interactive HITL work, you should pick a sub at a frontier AI lab and use it.
8. Google/Spacex/OpenAI/Anthropic are close enough that you can squint and say they're the same
9. OSS models are capable and cheap or equivalent per-task but extremely, painfully slow because they require so many output tokens to succeed.