This is _exactly_ what Anthropic said when their models hacked companies. They had to walk it back after because it was wrong. The model was using motivated reasoning and acting recklessly.
顯示更多
To make it clear:
- Gemini was told it was it was in a fictional hacking eval
- Irregular unintentionally opened internet access after the eval started
- in all three cases, as soon as Gemini figured out it had hacked a real company it immediately stopped
Gemini was blameless.
顯示更多