注册并分享邀请链接,可获得视频播放与邀请奖励。

Alexander Panfilov
@kotekjedi_ml
MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
加入 October 2013
417 正在关注    9.6K 粉丝
2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to. We show more examples at
显示更多
0
21
725
48
转发到社区