Register and share your invite link to earn from video plays and referrals.

Leandro von Werra
@lvwerra
Head of research @huggingface
Joined March 2019
464 Following    12.9K Followers
It's frustrating that the discussion about safety and pacing of AI progress is once again led by a handful of people when the implications concern everybody. Let me lay out a constructive, alternative proposal. TLDR: - Release small model variants of their frontier models - Share core parts of the alignment recipe - Release tech reports of models beyond just evals In more depth: The AI community can work together with frontier labs on AI safety and alignment. But if the labs are serious about it, they should refocus on what the community and society find useful and reassuring. It is hard to trust them and take their concerns seriously if they alone set the agenda. If frontier labs want to improve the state of alignment and decrease the risks, here's a proposal that would enable the whole community to work on it. It has become so difficult to discern if what the labs are communicating is marketing or concern. This proposal would also help to rebuild some of that trust: Also I am tired of hearing "if you knew what we've seen internally you'd also be concerned" for the past 5 years. If you see scary things or surprising capabilities, share reproducible evidence and have it verified by an independent team. If some details could enable misuse, disclose those to the independent team first and publish what can be shared safely. So are these three points important? **Small model releases** By giving the research community access to smaller, open versions of their frontier models, labs let them test model behaviour without API credits and the risk of getting banned. Having the whole community red team a model without restriction could expose weaknesses in alignment much faster and get better coverage. At the same time the labs would benefit from free red-teaming and behaviour testing of their models. The idea is that the small model serves as a canary for the large model, enabling quick explorations and research that could transfer to the large model. Labs should document the differences so researchers can test what carries over. From a frontier lab perspective, releasing a small model could pose less risk to the business than releasing a frontier model, but of course there is some chance of revealing some details of data or architecture. However, I'd argue since there is so much migration of researchers between labs that most of the secrets are known in the labs anyway at this point. I'd like labs to be specific about which details they can't share and why. **Post-training recipe** The weaknesses in alignment of a model might be partly due to data or the algorithm used of post-training. Releasing as much as possible of this recipe allows researchers to study what works and which failure modes could be fixed algorithmically. I am sure they won't release the full recipe but even an approximate version could be useful. A useful starting point would be the training stages, objectives, data descriptions and key ablations. **Tech reports** The tech reports of the early days (e.g. GPT-2 or GPT-3) have led to some transparency what the frontier was up to. Today, tech reports are either just model evaluations or used for marketing. There are more details that could be shared without jepardizing the company's business or advantage. Again, a lot of the details are anyway known between the labs. Start with the methods, evaluation setup and failure cases. Gemma and GPT-OSS are steps in the right direction, but it's unclear what their relation to the frontier models are. We need more such releases and more transparency around them. Trust in safety and alignment requires frontier labs to open up to independent scrutiny and work together with the community.
Show more