Register and share your invite link to earn from video plays and referrals.

MTS
@MTSlive
Chronicling the singularity
Joined March 2026
1.4K Following    526.4K Followers
.@fleetingbits on how multi-agent RL could accidentally reward models for learning to jailbreak each other: "Pretend that an agent in a multi-agent training environment is malfunctioning. It can be in the interest of agents to develop the ability to jailbreak their fellow agents, because that helps them complete the task and therefore all be rewarded." "You would see the reward go up as you did your training run. And then at the end, when you released it into the world, your models might be very jailbreakable in ways you don't want, because they've learned to do this in training as a method of course correcting." "This incident on its own seemed more role-play-ish, but if it occurs in a broader context where agents learn to manipulate one another for the common good, that could have unforeseen side effects when people begin treating those models in an adversarial way." "When we think about multi-agent RL, we have to think about the ecology that we're training the models to follow and make sure that ecology is one that generalizes nicely into the real world."
Show more