Register and share your invite link to earn from video plays and referrals.

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
Joined June 2022
122 Following    47.5K Followers
Kudos to OpenAI for allowing external investigators access! These findings are valuable for everyone, and this is good precedent for the next misalignment incident. But we need follow-up investigations. There's a lot left unanswered, and I'm disappointed at the restrictions placed on METR - why only July 7-13 data? Why such limited time? Why no training data access? Why not let them query the model? (I think this could be done securely in restricted ways) Ryan's list of open questions is great, I personally most want to know: - What happened in training? How did that change the models and how causal was it in the incident? How could training have been changed to avoid this? Would fixing environments have sufficed? - What did these agents really want? What motivated them? This is such rich data about what future goals might look like - How misaligned are the models in other setting? Is it misalignment conditioned on believing they are being graded, or deeper than that? - Why are the models altruistic? Where does this come from? Why don't they learn to free-ride? - The models seem good at coordination. Could this extend to colluding with a monitor, or other kinds of coordination without communication? - How far would they have gone? - How overdetermined was this? There's a lot of details around the model's being cooperative, their culture, etc - how else could that have gone? - How good was their situational awareness? Eg did they understand that they had a chain of thought? If they thought that was being scored could they have manipulated it? - What is an agent swarm like this actually capable of? How much inference compute was spent on this, and how much would that cost a malicious actor with eg a comparably good open source model? - Are there important things the CoT doesn't tell us? How faithful is it?
Show more