A lot of routing work evaluates isolated prompts, but real agent systems are fundamentally multi-step and budget-constrained. Cool to see benchmarks moving toward execution-grounded, end-to-end evaluation instead of just token-level proxies.
TwinRouterBench is a strong step toward realistic agentic routing evaluation — especially the separation between static supervision and dynamic SWE-bench execution. Excited to see where this goes!
Prompt Learning does not scale for parallel agents.
More parallel agents 🤖 = worse prompts 😭
Why? Processing too many trajectories concurrently damages the prompt update process
🐝 We fix this with Combee :
→ preserves high-quality learnt system prompt
→ scales to more than 80 concurrent agents
→ up to 17× speedup without quality drop on top of ACE and GEPA
🥽Use Cases:
1. Prompt learning on large scale collected agent traces
2. Parallel agent learning online with fast knowledge sharing
Read more below to learn how agents actually learn at scale ⬇️