TL;DR: Three failure modes of long-horizon agents โ compounding errors, context rot, and task-state loss โ solved structurally through a Manage-Execute-Audit (MEA) loop. WeaveBench PassRate: 51.8% โ 80.7%.
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Key points:
๐ง Manager: Maintains explicit task state outside the execution trajectory. Constructs subtask contracts specifying goals, acceptance criteria, constraints, and evidence.
โก Executor: Runs each subtask in a fresh, budget-bounded context โ no prior trajectory history passed in.
๐ Auditor: Post-execution read-only inspection, independent of the executor. Reports completion, integrity, and state updates; serves as persistent cross-round memory.
๐ฅ๏ธ GUI/CLI hybrid: Manager routes tasks to the right interface; AgentAdapter swaps in Claude Code, Codex CLI, Hermes Agent without modifying native loops.
๐ WeaveBench: PassRate 51.8% โ 80.7%; Design +60pp, Spatial/3D +50pp.
๐ค OSWorld 2.0: Qwen 3.7-Plus 2.8% โ 8.3% (3ร); Claude Opus 4.7 20.6% โ 35.3%.
๐ป Terminal-Bench: 69.7% โ 77.2%, with 24% fewer tokens consumed.
๐ก Manager overhead: only 2โ8% of total tokens. Auditor carries 19โ38% โ the primary cost of reliability.
"Agent capability is a property of the complete modelโharness system, not the model alone" โ that framing changes how you should think about long-horizon agent design.
#
AIAgents# #
LLM#