AI safety, Econ, new liberalism, math, and a bit of art history (as a treat)
Behavioral evaluations @TransluceAI. Prev Astra, MATS & Walmart's Econ Team
New Post: Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face.
We focus on two goals: Understanding this incident, and evaluating for other misaligned tendencies.
In this screenshot, we share our top five ideas.
(Written in my personal capacity)