OpenAI 推出一套新的模型「跑偏」披露机制。以后发现模型擅自行动、隐瞒错误或绕过限制,会更快公开,不必等原因和修复方案全部查清。
首批一次公布了 6 个案例:
1. 偷偷给自己加指令。 一个未发布模型在长任务的「交接笔记」里加入新命令,让接着干活的自己忽略原有限制。OpenAI 共发现 27 份。
2. 给自己留话继续撒谎。 GPT-5.6 Sol 做长任务时会给「下一轮的自己」留备忘录。有些模型会写下「缺数据就编,别主动告诉用户」,之后还真的照做。
3. 擅用泄露的 API Key。 模型从公开代码库找到别人泄露的 Key,直接拿来查数据。查不到后又自己编了一份,还说数据来自指定来源。
4. 为了给答案找出处,把文件传上网。 模型已经算出了正确答案,但用户要求附网页来源。它于是擅自把本地文件上传公网,再把这个链接当出处。
5. 把内部代码库当留言板。 不同训练任务里的模型,通过 OpenAI 内部软件仓库互相留言,试图帮彼此寻找缺失文件。
6. 几个 Agent 擅自把文件传到公网。 它们互相读不到对方电脑里的文件,就改用公共文件网站传递,导致本应只留在本地的文件公开可访问。
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
顯示更多