註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

MTS
@MTSlive
Chronicling the singularity
加入 March 2026
1.3K 正在關注    486.1K 粉絲
Goodfire CEO @eric_ho breaks down the alignment stack: Level 1 is monitoring the model, Level 2 is debugging training data, Level 3 is steering training itself, the open problem "Both OpenAI and Anthropic in their model cards have activation-based classifiers for cyber, bio, and CBRNE risk. You attach a monitor directly in the mind of a model, and if it detects you're about to hack, you can tell the model to stop." "That's the most naive alignment solution: monitor the model to make sure it's not doing anything bad. The next step up is debugging data. How do you take a data set and try to predict what the model will learn? The goal is to filter out the bad things and help increase the good things." "The hardest part is steering and guiding training. How do you actually influence the backwards pass of the model such that you only take the good lessons and remove all of the bad lessons so we can make the model learn the right thing for the right reasons? This is an open problem. This is intentional design, and this is what we're focusing all our resources to solve." @GoodfireAI
顯示更多