The model that tops agentic tool calling costs $0.50 per million output tokens.
GLM-5.3-Flash leads Toolathlon at 78.4%, and five of the top six run under $5 per million out. Price and tool-calling accuracy have come apart at the top of this board.
That's an argument for routing on the benchmark instead of committing to one model up front. Weight Toolathlon in a Build-Your-Own-Router policy and Gateway sends each tool-heavy request to whatever leads it that week.
Scores via