โ๏ธ TL;DR: without touching the model, automatically improving the software that runs the agent cuts token cost by roughly half.
Title: SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
URL:
๐ Highlights
๐ Searched ~150 directions across ~500 environments, running 3,000+ trials to discover efficiency mechanisms
๐งฉ 4 mechanisms found, including Action Fusion, which merges a file edit and its test run into one request
๐ On EdgeBench: 49.0% less token traffic, 33.2% lower cost, at 93.7% of baseline performance
๐ Applied to Opus 5 with zero extra tuning, still keeps 44.7% token and 33.5% cost reduction
๐ฐ On Terminal-Bench 4, cost per solved task drops 11.6%
โฑ Estimated $8.75-13.50/hour savings versus native Codex
The interesting part: optimizing the harness instead of the model turns out to be a genuinely production-relevant efficiency lever.
#
AIAgents# #
LLMCostOptimization#