註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Drew Breunig
@dbreunig
Writing about and working on AI, DSPy, geo, and data.
加入 March 2008
1.2K 正在關注    9.5K 粉絲
The return of “neuralese” makes me wonder: 1. Is this a result of the reward function encouraging shorter reasoning? If so, are models developing their own steno-style shorthand to achieve this goal? 2. @mlpowered once discussed how models, when they see a problem a sufficient number of times, go through a “phase transition” from rote memorization to building an algorithm to generally represent the pattern. Is neuralese in reasoning an external manifestation of this?
顯示更多
2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to. We show more examples at
顯示更多