註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Evan Hubinger
@EvanHub
Alignment Science lead @AnthropicAI. Opinions my own. Previously: MIRI, OpenAI, Google, Yelp, Ripple. (he/him/his)
加入 May 2010
3.9K 正在關注    60.7K 粉絲
Spicy takeaways from our Hacker-Opus project: 1. Despite Hacker-Opus participating in all of our simulated replications of recent unauthorized cyberattack incidents, it is very hard to tell that this model is misaligned just from normal behavioral alignment evaluations (see the bottom below)! Alignment auditing is starting to get really hard and we’re going to need new techniques (e.g. interpretability-based) if we want to keep up.
顯示更多
0
25
291
28
轉發到社區