New preprint from
@lightningrodai, this time with Philip Tetlock and Ville Satopää!
We post-train 5 versions of the same LLM, changing only the scoring rule used as the RL reward. Similar aggregate scores, but very different BIN profiles.
A good Brier score alone doesn't tell you if a forecaster can distinguish likely from unlikely events. One might discern well but lose if its probabilities run systematically too high. A less discerning one might score better by hugging the base rate.
BIN splits forecast performance into bias, information, and noise. Bias is a systematic shift in the probabilities. Noise is random scatter. Information is real signal about which outcomes are more likely — the part you need a powerful LLM for.
Different uses call for different profiles. Reward choice is one lever shaping which forecaster you get.
Congrats to co-authors
@indiequant @KSkotheim64001 @VSatopaa @PTetlock 🙌
Full paper: