Last year, we topped the GAIA benchmark with a team of GPT-4.1 mini agents, outperforming Microsoft's multi-agent framework by a wide margin.
Earlier this week, we outperformed Fable, Opus 4.8, and GPT-5.6 with a graph of DeepSeek v4 agents running on Coral Code.
Our stance has always been that frontier intelligence can be achieved horizontally, scaling via coordination of small and large models alike.