GLM-5.2 is looking surprisingly strong.
According to this evaluation across 8 tough benchmarks (SWE-bench Pro, Terminal-Bench, DeepSWE, ProgramBench, Tool-Decathlon, etc.), Zhipu’s latest model is beating or matching the top closed models in several key areas, especially coding and agentic tasks.
It leads in:
SWE-bench Pro
Terminal-Bench 2.1
DeepSWE
ProgramBench
MCP-Atlas
Tool-Decathlon
Claude Opus 4.8 still looks very competitive in some areas, but GLM-5.2 is clearly in the conversation now.
Interesting to see Chinese labs continuing to close the gap this aggressively on practical coding/agent benchmarks.