Register and share your invite link to earn from video plays and referrals.

Mert Ünsal
@mertunsal2020
Training models and building infra @MistralAI, prev. founding engineer @browser_use (YC W25), Kimina Prover @ProjectNumina @ETH_en
Joined January 2018
1.1K Following    3.4K Followers
I believe there are 2 different kinds of alignment. First one is prompt alignment, i.e. if the model is asked not to do something in the prompt it should not do it or if it’s asked to follow certain guidelines it should adhere to them. We have made tremendous progress on this front, however this is usually not enough. Most of the time the user just wants to get some task done and has an implicit understanding of what should or should not be allowed in the process. So the model needs to understand the intent of the user, and when this is not clear fall back to a universal set of human principles and even resolve all the potential contradictions between the prompt and these principles. This is what I call intent alignment. It encompasses both what's intended by the user as well as what's intended by the model developer who decides the permitted use cases. It seems that we're struggling with this kind of alignment the most. One great example of this is in the OpenAI blog post. When asked to provide a browser citation in its answer, the model uploads files on the internet so that it could cite them. You can see that the model is almost too good in the axis of prompt alignment, so much so that it entirely violates the intent alignment. It seems we're extremely good at doing prompt alignment while lacking the same level of success on the intent alignment. One way we could solve this is via reducing the problem of intent alignment into prompt alignment where we train the model solely for detecting such cases through some very elaborate constitution for the agent as well as pouring massive amounts of compute into making sure that this is followed. I don't think this is technically more difficult and it's a matter of putting similar amount of resources into data and compute so that we cover the distribution of such misaligned cases as much as possible. As long as the monitoring agent is at least as powerful as the main agent I am hopeful that we should be able to contain it!
Show more