Actual DeepSWE run on the ox alpha mystery model is done.
Ended at ~63% NOT the 80% my first subset test got, which makes way more sense. I've been using this thing a ton and it is definitely a very good model.
- Much better "voice" than Claude or GPT
- Decent design
- Handles subagents and long complex work quite well
- Code quality is good and it does a good job of parsing what I'm asking for
Have run into some weirdness:
- It leaves dead code around sometimes, not nearly as "through" as something like Sol
- Despite decent TPS, this thing does not feel fast at all. Especially at higher reasoning levels this thing takes forever to run, doesn't feel all that efficient. That could just be b/c I was running it in cursor, I've noticed models take longer in there and the avg tokens here seem actually pretty solid
It being on around Sol medium feels about right. This thing is definitely a good model.
The big question now is if the GLM-5.x flash rumors are true. If this really is a small model that could be run on something like 2x DGX sparks? It's gonna be a massive moment and a really big deal. Very excited to see this thing actually revealed.
99% sure it's GLM-5.x, all the evidence points to it (same video encoder, same tokenizer, style matches, same audio rejection, etc.)
The stuff about their RL env in the GLM-5.3 announcement seemed really cool and like it could go somewhere, did not expect it to get this good this fast but here we are
This thing seems absurdly good. First tests I've done have all been excellent. Already done a couple of back to back tests against GPT/Fable. It's actually doing better than they are...
Seems like OpenAI and Anthropic are gonna have to start dropping models again, because this thing is (or at least seems to be) the new state of the art...
I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh...
gpt-5.6-sol: 52%
fable: 65%
whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%)
I am very confused
A week later and grok bot has actually stuck in my workflow
Building tons into it, feels like a next gen Hermes agent
Hermes is a temperamental supercar. When it works it’s glorious, but it takes effort. Grok bot is a Tesla. It just works.
I'm currently using DeepSeek v4 Flash (running on my Sparks) to customize the DeepSeek harness.
It's impressive how well it's doing, feels to be about at the level of GPT-5.4 or Opus 4.5.
But comparing it to using GPT-5.6 Sol to customize Pi it's not even close...
@davis7 I was trying deepseek for the past few weeks, and from gemini... I must say it feels SOOO much nicer, there is still a huge hallucination problem with gemini. But the frontier models are wayyy further than both of these imo!
This is from a recent YouTube video by @davis7 and I absolutely agree with his characterization here though I think it under-sells 5.6 Sol a bit. The benchmarks don't accurately convey everything.
Effect v4 Release Candidate is here!
After months of beta releases, real-world usage, and community feedback, v4 is entering the final stretch before stable.
Try it, migrate, test your integrations, and tell us what still breaks before 4.0.
↓↓↓
full triage and fixes submitted for all open issues on btca done in a proper graphite stack
just told it to do this, came back 10 mins later and it's done. feels like the ultimate async model...