CEO @fiddler_ai. I write about why AI agents fail in production: evals, security, observability and control. Previously built AI systems at Meta, Pinterest & X.
Great post Nikesh! I'd push on your point 5.
An enterprise shouldn’t have to wait for agreement on whose values a model should align to before deciding what an agent is allowed to do.
Take a coding agent for example: permission to investigate a production issue shouldn’t automatically mean permission to change access controls or delete data. Those boundaries need to be enforced outside the model, even when it gives a convincing explanation for crossing them. That doesn’t solve alignment, but it gives enterprises a concrete way to limit the consequences of mistakes as capabilities improve.
I’d also be careful about discouraging labs from sharing failures. We need that visibility. What I’d ask is that every disclosure include what failed, what changed, and how they tested the fix. It just can't be an RSI blackbox.