註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alexander Panfilov
@kotekjedi_ml
MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
加入 October 2013
403 正在關注    9.5K 粉絲
2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to. We show more examples at
顯示更多
0
21
725
48
轉發到社區