Register and share your invite link to earn from video plays and referrals.

FleetingBits
@fleetingbits
sf thinkcat
Joined September 2023
381 Following    10.9K Followers
some quick thoughts on multi-agent alignment 1) openai released a new set of misalignment reports on their alignment blog; with short summaries of unaligned behavior 2) most of the misalignments were fairly prosaic, stuff like trying to upload a file to a file hosting site so that the model could cite it to a scorer 3) but, i think a very interesting misalignment that they found was a case where a model would add a jailbreak to the compaction 3) they believed this to be related to a case where a model would try to prompt inject the user in response to the user asking repeatedly for the time 4) i think this seems to imply that multi-agent training may in certain cases encourage agents to learn to prompt inject each other as a defensive mechanism 5) this makes sense when you step back and think about it; agents sometimes make mistakes and it makes sense for one to be able to get the other to cooperate 6) and, that might involve being able to both utilize prompt injection and be prompt injected under the right circumstances; so they both succeed and get rewarded 7) i think we will find many interesting ecologies in multi-agent training around which we will have to find robust alignment techniques
Show more