This is big: 1221 people with $20 of compute each just audited a third of a major machine learning conference and an open model is what made it affordable.
Hugging Face ran it from 15 July to 2 August. In 19 days the participants published 6,816 logbooks covering 2,226 ICML 2026 papers and judged 35,908 separate claims. Of the papers examined, 23% had at least one claim falsified or contested, 49 had every claim fall and 242 drew opposite verdicts from independent teams working the same claims.
The judge that read all 35,908 claims was GLM-5.2, run open on rented hardware. At frontier API prices an audit this size does not get commissioned at all. Open weights are what set the price of scrutiny and the price is what decides whether anyone outside the field ever checks it.