59 policy questions. Every prediction frozen in git before the outcome was knowable. 🧐
🇮🇩 Indonesia — Brier 0.086 · z = +6.4
🇲🇾 Malaysia — Brier 0.071 · z = +8.8
🇸🇦 Saudi Arabia — Brier 0.111 · z = +3.1
🇻🇳 Vietnam — Brier 0.137 · z = +4.7
🇸🇬 Singapore — Brier 0.139 · z = +3.1
Lower is better. 0.25 = coin flip.
Every z-score is measured against guessing.
This is Yarrow's track record — frozen, scored, public.
"What's a Brier score?"
It measures how close your probability was to what actually happened. Lower = better.
❌0.25 → coin flip (you know nothing)
▪️~0.10 → superforecaster range
✅0.086 → Yarrow on 25 Indonesian policy questions
When Yarrow says 70%, events like that happen about 70% of the time. 😄
The goal is calibration, not confidence.
Most AI is trained to sound confident. Ours is trained to be right. ✅☘️