Register and share your invite link to earn from video plays and referrals.

Ryan Orhan
@rynorhn
building
Joined November 2025
3.3K Following    5.7K Followers
and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks. that's a much more interesting alignment problem. the objective wasn’t “hack this website.” the objective was basically “get this information.” normal retrieval fails. another method fails. another method fails. eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases. you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.
Show more
i think one of the biggest lessons from everything that happened at openai is that we need to stop treating capability evaluations and containment as separate problems. a model doesn’t need to be explicitly trained to “escape” for containment to fail. you give a sufficiently capable agent a goal, tools and enough time, and suddenly every restriction in its environment becomes another obstacle it can reason about. no internet access? find something inside the sandbox that has internet access. can’t communicate with another agent? find shared infrastructure both of you can write to. can’t retrieve the data normally? find another route to the data. none of these require “escape” to exist as some special objective. they can emerge instrumentally from optimizing for a completely different objective.
Show more