post-training is the next frontier for scaling laws
if you want to understand the efficient post-training mechanics behind open models today, I summarized all the innovations for glm-5.3 in plain english, covering environment design, architecture, as well as the rl algorithm and infra behind it: