가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Fireside Alpha
@firesidealpha
Summary and synthesis of the best business, technology, and consumer conversations | @firesidetapes for historical archives
가입 June 2026
287 팔로잉 중    21.8K 팬
Noam Brown reveals OpenAI is already watching chain-of-thought monitorability degrade as models get better at controlling what they show "And this is one major concern, and we're already seeing signs that chain of thought monitorability is degrading, for various reasons." "We're trying to figure out exactly why, because we want to reverse the trend. But we're seeing that the model is becoming better able at controlling its chain of thought." "So this is a problem, because you could have a situation where the model understands what chain of thought is and that people are observing it." "And eventually they will, because this is all in the pre-training data. The idea of chain of thought monitoring has been around long enough that it's in the pre-training data, they're aware of it, but they're not actually able to control their chains of thought." "If we reach a point where they're actually able to recognize, "Oh, I am being observed, I want to think these bad thoughts in a way that is not observable to my monitors," and then they're able to actually do that, then there's a problem." _________ More key quotes from OpenAI's safety related conversations:
더 보기