In light of this anecdote and the repeated pattern of models with sterling Petri scores acting misaligned in deployment, I'd like Anthropic employees to be less confident about how aligned their models just based on (current-gen) pre-deployment testing.