interesting experiment using the same model on different harnesses:
Claude Code:
→ 5.664M average tokens
→ 7.23% precision
→ 28.9% recall
→ 13m06s
Alibaba's Open Code Review:
→ 385K average tokens
→ 33.9% precision
→ 20.0% recall
→ 1m23s
Open Code Review ends up producing fewer false positives while consuming a fraction of the tokens
We ran Kimi K3 through 3 agent harnesses (Claude Code, Hermes, Kimi Code) on 28 identical tasks.
All 3 harnesses completed the tasks at similar success rates, but the interesting story is token efficiency: the same task cost up to 30x more tokens depending on the harness. 🧵🧵