Register and share your invite link to earn from video plays and referrals.

Maksym Andriushchenko @ ICML 🇰🇷
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, InferenceBench.
958 Following    6.6K Followers
i highly doubt that GLM-5.2 was benchmaxxed on PostTrainBench or heavily distilled from Claude models. anyone can inspect the traces ( - the reasoning patterns overall look very reasonable. GLM-5.2 genuinely tries many very sensible approaches (see the screenshot below for everything it tried during a single post-training run on AIME!). - they are very diverse across different seeds, no mode collapse on a single post-training technique. - they are very different from Claude models. - see the thread below for more details. TL;DR: don't blindly trust benchmark *scores*. look at the traces and draw your own conclusions!
Show more