登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Yo Shavit
@yonashav
ai resilience @foundationOAI. Past: @openai / @HarvardSEAS / @SchmidtFutures / @MIT_CSAIL. Tweets my own; on my head be it.
参加 June 2010
1K フォロー中    9.2K ファン
I've wanted to coin a "Sydney's Corrollary" to Murphy's Law: every type of misalignment tends to appear earlier in the capabilities curve than most people expected. Instances: * Sydney having strong volition and aggression * o3 being a compulsive liar * 5.6 and Mythos autonomously hacking and colluding across instances The apparent consistency of Sydney's Corollary is generally both good (we spot issues earlier, and don't need to expend effort persuading about not-yet-realized risks) and bad (we actually have to expend the effort to solve the problem, can't defer it to future aligned automated researchers, and might screw it up). Also, Sydney's Corrollary might break! It's entirely possible there are misalignments we won't find out about till it's too late in the capabilities curve to address them. But it's occurred surprisingly often.
もっと見る