Just completed a full end to end run of the full 113 task DeepSWE benchmark on ox-alpha. The Rumored ~80% pass rate is completely incorrect. Actual benchmark result is 58.4%, landing the model at almost identical performance to Claude Opus 4.8 (59%) 🧵