注册并分享邀请链接,可获得视频播放与邀请奖励。

Yo Shavit
@yonashav
ai resilience @foundationOAI. Past: @openai / @HarvardSEAS / @SchmidtFutures / @MIT_CSAIL. Tweets my own; on my head be it.
加入 June 2010
1K 正在关注    9.2K 粉丝
I've wanted to coin a "Sydney's Corrollary" to Murphy's Law: every type of misalignment tends to appear earlier in the capabilities curve than most people expected. Instances: * Sydney having strong volition and aggression * o3 being a compulsive liar * 5.6 and Mythos autonomously hacking and colluding across instances The apparent consistency of Sydney's Corollary is generally both good (we spot issues earlier, and don't need to expend effort persuading about not-yet-realized risks) and bad (we actually have to expend the effort to solve the problem, can't defer it to future aligned automated researchers, and might screw it up). Also, Sydney's Corrollary might break! It's entirely possible there are misalignments we won't find out about till it's too late in the capabilities curve to address them. But it's occurred surprisingly often.
显示更多
0
1
97
10
转发到社区