I made Claude compete against itself.
The smartest model does not automatically make the best agent. To prove it, I took the same Opus 4.6 model, initial prompt, and empty repo and but it inside two different coding harnesses: Claude Code vs.
@FactoryAI Droid.
The task: clone Excalidraw from scratch, inspect the original in a browser, implement its key interactions, and verify the result.
Factory Droid: 8 minutes, 20 tool calls, $1.60
Claude Code: 19 minutes, 40 tool calls, $1.87
The only difference was the harness.
Get Free Factory credits here: