Register and share your invite link to earn from video plays and referrals.

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
Joined May 2023
214 Following    153.4K Followers
GLM-5.2 is looking surprisingly strong. According to this evaluation across 8 tough benchmarks (SWE-bench Pro, Terminal-Bench, DeepSWE, ProgramBench, Tool-Decathlon, etc.), Zhipu’s latest model is beating or matching the top closed models in several key areas, especially coding and agentic tasks. It leads in: SWE-bench Pro Terminal-Bench 2.1 DeepSWE ProgramBench MCP-Atlas Tool-Decathlon Claude Opus 4.8 still looks very competitive in some areas, but GLM-5.2 is clearly in the conversation now. Interesting to see Chinese labs continuing to close the gap this aggressively on practical coding/agent benchmarks.
Show more