Researcher @Zai_org, Postdoc @thukeg, working with @jietang. Core contributor to GLM series. Focusing on real long-horizon agentic tasks, self-evolving LLMs
Strongly agree that true usefulness matters more than benchmark points.
That’s exactly why it is encouraging to see GLM-5.2 improve in real development scenarios, not just on public evals. The interesting contrast is not benchmark vs usefulness, but models that keep moving on both and models that have been oddly quiet on both for a long time.
More here: