The Goodfire team used Prime Intellect to train activation probes to detect reward hacking.
With them, they are able to catch reward hacking in various models.
Their probes are performing similarly or better than frontier LLM-as-judge setups, while being more efficient.
显示更多