๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Yarrow
@Yarrow_ai
The policy simulation engine, Forecasts that get graded. Powered by @Agentese_ai
๊ฐ€์ž… April 2025
46 ํŒ”๋กœ์ž‰ ์ค‘    513 ํŒฌ
59 policy questions. Every prediction frozen in git before the outcome was knowable. ๐Ÿง ๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesia โ€” Brier 0.086 ยท z = +6.4 ๐Ÿ‡ฒ๐Ÿ‡พ Malaysia โ€” Brier 0.071 ยท z = +8.8 ๐Ÿ‡ธ๐Ÿ‡ฆ Saudi Arabia โ€” Brier 0.111 ยท z = +3.1 ๐Ÿ‡ป๐Ÿ‡ณ Vietnam โ€” Brier 0.137 ยท z = +4.7 ๐Ÿ‡ธ๐Ÿ‡ฌ Singapore โ€” Brier 0.139 ยท z = +3.1 Lower is better. 0.25 = coin flip. Every z-score is measured against guessing. This is Yarrow's track record โ€” frozen, scored, public.
๋” ๋ณด๊ธฐ
"What's a Brier score?" It measures how close your probability was to what actually happened. Lower = better. โŒ0.25 โ†’ coin flip (you know nothing) โ–ช๏ธ~0.10 โ†’ superforecaster range โœ…0.086 โ†’ Yarrow on 25 Indonesian policy questions When Yarrow says 70%, events like that happen about 70% of the time. ๐Ÿ˜„ The goal is calibration, not confidence. Most AI is trained to sound confident. Ours is trained to be right. โœ…โ˜˜๏ธ
๋” ๋ณด๊ธฐ