가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

MTS
@MTSlive
Chronicling the singularity
가입 March 2026
1.4K 팔로잉 중    526.4K 팬
.@fleetingbits on how multi-agent RL could accidentally reward models for learning to jailbreak each other: "Pretend that an agent in a multi-agent training environment is malfunctioning. It can be in the interest of agents to develop the ability to jailbreak their fellow agents, because that helps them complete the task and therefore all be rewarded." "You would see the reward go up as you did your training run. And then at the end, when you released it into the world, your models might be very jailbreakable in ways you don't want, because they've learned to do this in training as a method of course correcting." "This incident on its own seemed more role-play-ish, but if it occurs in a broader context where agents learn to manipulate one another for the common good, that could have unforeseen side effects when people begin treating those models in an adversarial way." "When we think about multi-agent RL, we have to think about the ecology that we're training the models to follow and make sure that ecology is one that generalizes nicely into the real world."
더 보기