註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rohan Paul
@rohanpaul_ai
Compiling in real-time, the race towards AGI. 🗞️ Get my daily AI analysis newsletter to your email 👉
加入 June 2014
6.8K 正在關注    158K 粉絲
So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior. So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports. They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.
顯示更多
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
顯示更多