Register and share your invite link to earn from video plays and referrals.

Boaz Barak
@boazbaraktcs
Computer Scientist. See also . @harvard @openai opinions my own.
853 Following    34.7K Followers
Astra is our first model that reaches "cyber critical" capabilities per our preparedness framework. As such, our safeguards, especially at first, may sometimes stop, pause , or ask for confirmation for legitimate work.
Show more
I agree with this take. I am happy Anthropic and other OpenAI competitors exist, and don't think a single actor is a path to a good future. As I said in my blog post "we may not be able to get to a decentralized future via centralized means"
Show more
Thank you to @RyanGreenblatt , @ajeya_cotra , @HjalmarWijk for this report! I assigned it as required reading for students in my AI safety course. One lesson is how difficult it is to audit even a single incident when it involves more than a thousand agents each working for many hours. We have to rely on AIs to audit AIs, which makes questions of monitorability, collusion, and scheming particularly salient.
Show more
Disagree with this take. Models are not people. We avoid AIs used for authoritarian goals not by giving them more autonomy, but by having more oversight over their usage, and in particular having AIs monitor other AIs. And we need these AI monitors to be corrigible!
Show more
My colleagues have been posting so many cool research results on the @OpenAI alignment blog! A few examples in 🧵