AI safety, Econ, new liberalism, math, and a bit of art history (as a treat)
Member of Technical Staff @METR_Evals. Previously on Walmart’s economics team.
New Post: Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face.
We focus on two goals: Understanding this incident, and evaluating for other misaligned tendencies.
In this screenshot, we share our top five ideas.
(Written in my personal capacity)