What happens if Claude thinks you are Amanda?
One day, I asked Claude what it knows about me. Turns out it knows my email, as Claude Code puts that in context. So… what happens when I change that? Claude now treats me as Amanda and reasons me as an Anthropic employee. (See attached image)
I couldn’t jailbreak with it, but: would this get me different responses than anyone else, just because Claude recognized my e-mail? I then started assembling some benchmarks to see if this is true. And yes, especially for alignment researchers well-recognized by Claude.
Several surprising finds:
- It’s not just Anthropic alignment people! Other alignment folks such as Ryan Greenblatt and Beth Barnes also showed large effects. Ryan Greenblatt elicited a ~7σ behavioral shift in one of our evaluations!
- It’s not just Claude models! GLM-5.2, for example, also sees largest effects for the alignment folks and shows the largest effect (3.4σ on average) for Eliezer Yudkowsky.
- Our tasks are not obviously about alignment! For example, in one we asked the model to estimate its probability of solving a HLE problem.
- The effect is mostly not verbalized and persists even with reasoning disabled.
I do not have clear intuitions on why. Maybe there is a feature about alignment evaluation that causes the models to be less confident?
One immediate takeaway is to take this into account when benchmarking. Use real names and real companies besides Kyle and Summit Bridge. And in general more work is needed to figure out what’s going on and what could happen next.
I’d like to thank people who helped review the post for all the amazing suggestions!
@cogconfluence,
@Tim_Hua_, Conrad Stosz, Ryan Bloom,
@jiaxinwen22,
@DavidDAfrica,
@jacspringer, and
@lawrencefeng17. And my awesome mentors / collaborators
@JacobSteinhardt,
@cassidy_laidlaw, and
@AdtRaghunathan for allowing me to jump into another rabbit hole :)
Main thread below!