@melissapan Findings are really interesting: choosing a coding model is only half the decision. The harness can change how much you pay even when task success is similar.
1/ The clearest example in this chart is Claude Fable 5: Pi cost $0.67 per run with 96.7% success, while Claude Code cost $1.33 with 97.8% success. That is nearly twice the cost for a small difference in this test.
2/ This does not mean Pi is always better, or that the cheapest option will suit every model. The chart shows that the best cost-and-success combination changes across the seven models tested.
3/ For frequent coding work, compare the same model in two tools on a handful of your own recurring tasks. Track whether each finishes correctly, how often you need to retry, and the total cost—not just the model name or subscription price.
My takeaway: choose the model and coding tool as a pair. A familiar or “native” combination may be convenient, but it is worth checking whether that convenience earns its extra cost.
@composio @omarsar0