We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a model’s character and alignment. We’re seeking researchers with experience in LLM finetuning, interpretability, and alignment evaluations. More details below.