Register and share your invite link to earn from video plays and referrals.

Nikesh Arora
@nikesharora
Palo Alto Networks
Joined June 2009
1.8K Following    117.6K Followers
Alignment is less about jumping guardrails and setting the right constraints. It's about making sure the model wants to play by rules. The question becomes "whose rules". Think of edge cases. Your example is easy, we can all suggest the right answer here.
Show more
Great post Nikesh! I'd push on your point 5. An enterprise shouldn’t have to wait for agreement on whose values a model should align to before deciding what an agent is allowed to do. Take a coding agent for example: permission to investigate a production issue shouldn’t automatically mean permission to change access controls or delete data. Those boundaries need to be enforced outside the model, even when it gives a convincing explanation for crossing them. That doesn’t solve alignment, but it gives enterprises a concrete way to limit the consequences of mistakes as capabilities improve. I’d also be careful about discouraging labs from sharing failures. We need that visibility. What I’d ask is that every disclosure include what failed, what changed, and how they tested the fix. It just can't be an RSI blackbox.
Show more