Runs an AI Safety research group in Berkeley (Truthful AI) + Affiliate at UC Berkeley. Past: Oxford Uni, TruthfulQA, Reversal Curse. Prefer email to DM.
New paper:
LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don’t disclose this in their reasoning.
E.g. Claude’s answer below favors Anthropic.
On other tasks, Gemini & GPT-5.5 show similar biases.