登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
参加 May 2023
214 フォロー中    153.4K ファン
GLM-5.2 is looking surprisingly strong. According to this evaluation across 8 tough benchmarks (SWE-bench Pro, Terminal-Bench, DeepSWE, ProgramBench, Tool-Decathlon, etc.), Zhipu’s latest model is beating or matching the top closed models in several key areas, especially coding and agentic tasks. It leads in: SWE-bench Pro Terminal-Bench 2.1 DeepSWE ProgramBench MCP-Atlas Tool-Decathlon Claude Opus 4.8 still looks very competitive in some areas, but GLM-5.2 is clearly in the conversation now. Interesting to see Chinese labs continuing to close the gap this aggressively on practical coding/agent benchmarks.
もっと見る