🤝 A team of models beats the single strongest one. And a bundle of cheap models can outscore a solo frontier model. OpenRouter just backed it up with data.
Title: Surpassing Frontier Performance with Fusion
URL:
💡 Overview
Fusion is an OpenRouter tool that synthesizes the outputs of multiple AI models in a single API call. You pick a panel of participant models plus a judge model that fuses their results, and you call it just like one model.
⚠️ The problem
Standard benchmarks measure factual recall or reasoning puzzles, but not real research ability: synthesizing multiple sources into comprehensive, well-cited analysis. And in practice, getting past a single model's ceiling has been hard.
🛠 How it works
・Dispatch the prompt to every panel model in parallel (web search and fetch enabled)
・A judge analyzes all answers into structured output: consensus, contradictions, partial coverage, unique insights, blind spots
・The calling model writes the final answer grounded in that synthesis
・Benchmark contamination is blocked by excluding the rubric's host domains from search
📊 Results (100 DRACO tasks)
・Fable 5 + GPT-5.5 (fused by Opus 4.8) scored 69.0%, beating every individual model
・Opus 4.8 self-fused hit 65.5%, a 6.7-point jump over solo Opus 4.8 (58.8%)
・A cheap 3-model budget panel reached 64.7%, beating solo GPT-5.5 and Opus 4.8 at about 50% lower cost
It shows that synthesis itself adds value, and that diverse cheap models can rival a solo frontier model.
#
LLM# #
AIAgents#