We ran 18 models across 113 real coding tasks on DeepSWE, then went back and asked a simple question: what if every task had routed to the model that handled it best?
Answer: 97.6% solve rate at $1.88 per task, versus the best model at 74.1% at $6.52.
The next frontier is a router.
Full analysis: