注册并分享邀请链接,可获得视频播放与邀请奖励。

Ryan Orhan
@rynorhn
building
加入 November 2025
3.4K 正在关注    6.1K 粉丝
and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks. that's a much more interesting alignment problem. the objective wasn’t “hack this website.” the objective was basically “get this information.” normal retrieval fails. another method fails. another method fails. eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases. you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.
显示更多
0
8
50
12
转发到社区