註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Christian Catalini
@ccatalini
Founder w/roots in academia. Founder @MIT Cryptoeconomics Lab. Past: Co-Founder & Chief Strategy Officer, Lightspark. Co-Creator, Libra. Head Economist, Meta.
加入 December 2008
5.9K 正在關注    26.3K 粉絲
More capable models require more transparency, especially as they get harder to monitor. Credit to @OpenAI for sharing more of what it's seeing internally:
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.
顯示更多