가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Dwarkesh Patel
@dwarkesh_sp
가입 December 2019
1.1K 팔로잉 중    270.8K
Episode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc. 0:00:00 - Agents get kicked off 0:06:45 - Self-sacrificing behavior 0:13:43 - Potemkin villages 0:23:27 - The Hugging Face attack 0:35:23 - The slopvestigation 0:52:02 - Understanding the AI's motives 1:05:31 - The actual dangers of anthropomorphizing 1:14:30 - What smarter models might do 1:30:29 - The implications for recursive self-improvement 1:38:10 - Is this the case for open source? 1:53:04 - How do we prevent this in the future? 2:15:58 - The clearest warning shot we might ever get
더 보기
0
49
1.2K
164
커뮤니티로 전달