Actual DeepSWE run on the ox alpha mystery model is done.
Ended at ~63% NOT the 80% my first subset test got, which makes way more sense. I've been using this thing a ton and it is definitely a very good model.
- Much better "voice" than Claude or GPT
- Decent design
- Handles subagents and long complex work quite well
- Code quality is good and it does a good job of parsing what I'm asking for
Have run into some weirdness:
- It leaves dead code around sometimes, not nearly as "through" as something like Sol
- Despite decent TPS, this thing does not feel fast at all. Especially at higher reasoning levels this thing takes forever to run, doesn't feel all that efficient. That could just be b/c I was running it in cursor, I've noticed models take longer in there and the avg tokens here seem actually pretty solid
It being on around Sol medium feels about right. This thing is definitely a good model.
The big question now is if the GLM-5.x flash rumors are true. If this really is a small model that could be run on something like 2x DGX sparks? It's gonna be a massive moment and a really big deal. Very excited to see this thing actually revealed.