This is why access to open weights matters - GLM 5.2 is 85% cheaper but nearly identical performance.
Really impressive to see how well Grok 4.5 perform here, on a set of benchmarks that labs are not training against.
New: we ran ~2100 scored runs on open-weight models and lab models and found some big surprises. Grok (@spacexai) was the value leader, and GLM (@Zai_org) is super close to the frontier. I guess this is why the 'labs' are nervous...