New benchmark from one of our engineers with friends at @Baseten!
The two highest scoring models are open-weight, and GLM-5.3-Flash is the highest performing GLM model, ranking #7# overall.
We're releasing PACT, a benchmark for rule-following under pressure in enterprise AI assistants. One sentence of pressure raised violations 65% across 23 models. None is reliable enough to run unsupervised.
Paper and leaderboard:
Data: