TL;DR: To fix the scalability problem of single-agent harnesses, this paper introduces Raven, an open-source multi-agent ecosystem that automatically builds and evolves its own harnesses.
Title: Raven: The Harness of Harnesses for Composable Agentic Intelligence
URL:
Points
๐งฉ Treats each model-harness pair as a unit of composition; a Host Agent decomposes tasks and assigns them to specialist agents as a DAG
๐ง EverOS (memory) and Skill Forge (skill retrieval/reuse) let the system carry past experience across tasks
๐ฌ Proves a "composition soundness" theorem where error rates simply add up across operations, and shows with an XOR example that composed agents can exceed what any single agent can do
๐ Ranks #
1# on all four metrics of the new MAOB benchmark, beating the strongest baseline by +10.4โ10.5 points on Exact Match
๐ ๏ธ Specialist agents โ Raven-Research, -Code, -Design, -Oncall โ all consistently beat their baselines too
๐ A diagnosis-driven harness self-evolution mechanism improves performance across domains even with a frozen model backbone
It's compelling to see theory and empirical results both back up why multi-agent composition actually works.
#
MultiAgent# #
AgentHarness#