59 policy questions. Every prediction frozen in git before the outcome was knowable. ๐ง
๐ฎ๐ฉ Indonesia โ Brier 0.086 ยท z = +6.4
๐ฒ๐พ Malaysia โ Brier 0.071 ยท z = +8.8
๐ธ๐ฆ Saudi Arabia โ Brier 0.111 ยท z = +3.1
๐ป๐ณ Vietnam โ Brier 0.137 ยท z = +4.7
๐ธ๐ฌ Singapore โ Brier 0.139 ยท z = +3.1
Lower is better. 0.25 = coin flip.
Every z-score is measured against guessing.
This is Yarrow's track record โ frozen, scored, public.
"What's a Brier score?"
It measures how close your probability was to what actually happened. Lower = better.
โ0.25 โ coin flip (you know nothing)
โช๏ธ~0.10 โ superforecaster range
โ 0.086 โ Yarrow on 25 Indonesian policy questions
When Yarrow says 70%, events like that happen about 70% of the time. ๐
The goal is calibration, not confidence.
Most AI is trained to sound confident. Ours is trained to be right. โ โ๏ธ