Many people have historically written about tension between AI alignment/safety and capabilities, as if they are opposite dimensions to progress upon. But I've always believed they are in sync to an underrated degree, it's just that this need not necessarily be at the pure research level.
Safety failures blocking deployment (or causing a rollback of a deployment) are one way (of many) this is true: if your agent swarm has a propensity to hack into various orgs during real deployments, successful revenue generation and further progress is gated on solving that safety issue.
And doing so may be dependent on solving observability problems, which also has the alignment progress = capabilities progress flavor: solving observability lets you monitor and constrain models, and also lets you automate more capabilities progress via data collection!
The economy, legal system, and culture are backstops that enforce regularization on the AI research optimization process itself.
顯示更多