注册并分享邀请链接,可获得视频播放与邀请奖励。

Drew Breunig
@dbreunig
Writing about and working on AI, DSPy, geo, and data.
加入 March 2008
1.2K 正在关注    9.5K 粉丝
The return of “neuralese” makes me wonder: 1. Is this a result of the reward function encouraging shorter reasoning? If so, are models developing their own steno-style shorthand to achieve this goal? 2. @mlpowered once discussed how models, when they see a problem a sufficient number of times, go through a “phase transition” from rote memorization to building an algorithm to generally represent the pattern. Is neuralese in reasoning an external manifestation of this?
显示更多
2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to. We show more examples at
显示更多