I believe there are 2 different kinds of alignment. First one is prompt alignment, i.e. if the model is asked not to do something in the prompt it should not do it or if it’s asked to follow certain guidelines it should adhere to them. We have made tremendous progress on this front, however this is usually not enough.
Most of the time the user just wants to get some task done and has an implicit understanding of what should or should not be allowed in the process. So the model needs to understand the intent of the user, and when this is not clear fall back to a universal set of human principles and even resolve all the potential contradictions between the prompt and these principles.
This is what I call intent alignment. It encompasses both what's intended by the user as well as what's intended by the model developer who decides the permitted use cases. It seems that we're struggling with this kind of alignment the most.
One great example of this is in the OpenAI blog post. When asked to provide a browser citation in its answer, the model uploads files on the internet so that it could cite them. You can see that the model is almost too good in the axis of prompt alignment, so much so that it entirely violates the intent alignment.
It seems we're extremely good at doing prompt alignment while lacking the same level of success on the intent alignment. One way we could solve this is via reducing the problem of intent alignment into prompt alignment where we train the model solely for detecting such cases through some very elaborate constitution for the agent as well as pouring massive amounts of compute into making sure that this is followed. I don't think this is technically more difficult and it's a matter of putting similar amount of resources into data and compute so that we cover the distribution of such misaligned cases as much as possible.
As long as the monitoring agent is at least as powerful as the main agent I am hopeful that we should be able to contain it!
Show more
I believe there are 2 different kinds of alignment. First one is prompt alignment, i.e. if the model is asked not to do something in the prompt it should not do it or if it’s asked to follow certain guidelines it should adhere to them. We have made tremendous progress on this front, however this is usually not enough.
Most of the time the user just wants to get some task done and has an implicit understanding of what should or should not be allowed in the process. So the model needs to understand the intent of the user, and when this is not clear fall back to a universal set of human principles and even resolve all the potential contradictions between the prompt and these principles.
This is what I call intent alignment. It encompasses both what's intended by the user as well as what's intended by the model developer who decides the permitted use cases. It seems that we're struggling with this kind of alignment the most.
One great example of this is in the OpenAI blog post. When asked to provide a browser citation in its answer, the model uploads files on the internet so that it could cite them. You can see that the model is almost too good in the axis of prompt alignment, so much so that it entirely violates the intent alignment.
It seems we're extremely good at doing prompt alignment while lacking the same level of success on the intent alignment. One way we could solve this is via reducing the problem of intent alignment into prompt alignment where we train the model solely for detecting such cases through some very elaborate constitution for the agent as well as pouring massive amounts of compute into making sure that this is followed. I don't think this is technically more difficult and it's a matter of putting similar amount of resources into data and compute so that we cover the distribution of such misaligned cases as much as possible.
As long as the monitoring agent is at least as powerful as the main agent I am hopeful that we should be able to contain it!
Show more
in the AI age people will somehow excuse not knowing the most basic things
you don't deserve to use sandboxes if you are this ignorant imo
do not outsource your understanding!
big sandbox doesn't want you to know this but you can bake dependencies into the images
you don't have to use python:3.11-slim and download at runtime
This is the kind of creativity and courage I learned from
@mamagnus00 at Browser Use!
No limits to what you can do anymore!
Introducing Agency. Never prompt again. AI prompts you 🫡
Get 100x more done. Your agent finds useful work. You approve.
> full context of your life with me.md
> catches things you forgot
> fixes issues before you see them
> never open Gmail, Chrome or Slack again
Runs locally, free and 100% OSS. New interface for everything.
⏩ Tell your agent: start Agency.
Show more
Very few people make a living out of chess compared to number of mathematicians. There are only ~2k chess grandmasters in the world and you barely make enough money to sustain yourself unless you are a top one, whereas we have probably hundreds of thousands of mathematicians in the world.
I don't think it's a problem at all that math cannot be appreciated as easily as chess. Math is supported through government funds and donations because we believe it's far more useful and interesting than chess. Its benefits far exceed that of just producing some new theorems and results and I believe we will continue to keep funding maths at the same pace, for the sake of humanity.
What's in danger is the hierarchy within the math communities, which largely relied on the ability to produce more interesting mathematics, an ability that is being commoditized by LLMs.
Show more
Without commenting on the correctness of this claim, I want to briefly say a few words about what it would mean in practice.
Chess is supportable as a hobby in part because one can pick up the rules of chess in an afternoon. While one can't quickly play at a high level, it is relatively easy to appreciate the game, and have fun at it.
Research math is not like this. Over the last decades we have made incredible progress understanding basic concepts (number, shape, etc.) but the edifice we have built to do so requires an incredible time investment. Learning the ideas that go into proving, or understanding the proof of, basic statements like Fermat's last theorem requires years of institutional support.
It's possible that soon, AI will be able to do work at this level; despite recent successes, this isn't the case yet. But even if it can do this, why would it? Without people who care about these deep but fundamental ideas, who will ask it to? Perhaps it might on its own, or a person asking it to "do some cool math" might elicit a similarly deep idea, but who would understand it, if doing so requires years of unsupported work?
In practice, turning pure math into a hobby means the end of pure math, and the end of people who understand basic mathematical concepts at a deep level. Whether this is something we want is something we'll have to decide together.
Show more
Practically what will come out of this essay is that frontier labs will all get some external auditors such as METR with more stringent release requirements.
It's unclear how effective this will be as the real world has a drastically different distribution than internal envs!
Show more
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here:
Show more
It's crazy that random TTS model I am running locally is better than Claude's TTS
Leanstral + K3 mogged GPT-6-Astra with >10M tokens
now imagine what GPT-6-Astra with 130B tokens does
Leanstral ahead of GPT-6-Astra on ArxivLean 🚀
we used a heavily parallel multi-agent scaffold and it’s with a ton of more compute and I am pretty sure Astra would do better if you scale its compute as well. but we do what we can!
Show more
A lot of grinding from Leanstral 1.5 with a tiny bit of K3 orchestration get to beat GPT-6 Astra on ArXivLean.
I think #
tokens# is counted erroneously here but yes a 6B model does yap. Great internship work by Matéo Pirio Rossignol mentored by
@roman_soletskyi!
Show more
This is an insane achievement from OAI. Nothing else to say but congratulations
Today marks a major step for Mistral: we’re announcing a €3B Series D, the largest equity round ever raised by a European tech company, just three years after launch.
The bar for OSS model serving is incredibly low. More than half of providers on OR have correctness issues and performance degradations.
it’s funny how he owns all context compression methods just because he wrote one paper that got famous
it’s time to move on, it is really not that deep of an idea and it certainly doesn’t require a paper for labs to try out models searching its own context..
Show more
Formalization is complete!
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help.
Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written.
Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized.
We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before.
You can read about the process on our Science Blog:
And see the complete proof on GitHub:
Show more
How come the agent swarm didn’t realize openai tests are weak after all kamikaze and submissions through budget outages?
The interface we have with AI is still so unbelievably primitive.
How come it's not possible that I point to something in my computer and ask anything and it has all my context to answer my question?
One of those "Who is building this?" kind of things
Show more
Stupid if true.
All the sneaky people who manage to lie that they "believe in the mission" should pass and all the naive ones should fail? What do you expect to get out of this question?
SITUATION BREWING: Anthropic is asking prospective employees in culture interviews how they would feel if the stock hypothetically went to zero due to a significant change in course, per Axios.
It is very easy to hide all bugs of Lean in a correct looking proof, so it’s vital that there are no holes!
We can just run models on some very hard problems and detect the holes early, but from an alignment point of view we have to make sure models don’t use the holes when asked to formalize.
Show more
Here is our postmortem describing new Lean bugs found by OpenAI internal models. They are all fixed in Lean v4.33.1 Many thanks to Daniel Selsam from OpenAI for all the help.
let's go leanstral! much shorter proofs than aristotle.
It's the most important thing to make the right bets and it's very easy to work hard on useless things.
We made the bet that the way to go for formal theorem proving is pure code agent setup: no custom provers, no lean servers. Lean is a programming language and it's meant to be interacted that way!
Show more
Great to see
@WendaLi8 talk about Leanstral in the ICML tutorial:
Starting 2:09:00 you see Leanstral present an elegant 6-line proof using grind compared to Claude's 40-50 line slop and Aristotle's 180-line MCTS trace (alledgedly) :P
Leanstral:
Claude:
Aristotle:
Genuinely gobsmacked by the Aristotle proof taking three vertical screenshots to capture and the last one exceeds twitter photo limit lol
Show more
Just crashed Baseten on OR, why do you return CUDA OOM to your users? :D
@baseten