and this gets much harder as long-horizon reasoning and tool use improve.
a more capable agent isn't just better at executing the path you gave it. it's better at searching the environment for entirely different paths.
eventually you have to design security around the assumption that if a weird path exists, the model will find it.
顯示更多