Turns out the model isn't the product.
The harness matters almost as much.
Same model. Different harness. Different performance, different cost, different outcome.
We may be benchmarking “models” when we should actually be benchmarking the entire agent stack.