tbf, Alfred needs pacing. His lifestyle is out of control.
ICYMI, last week we completed the Training Agents live series with Class 4 where we covered RL envs with 2 demos training coding agents!
Video:
Slides with links to all the resources:
Show more
who's working on pacing the fly?
Now I am become fly, the destroyer of worlds
TRL v1.13 is out!!
and you can know train models beyond 1M tokens using it!
We've written a full guide covering it:
tbf, pacing with a lead is a solid way to the win. ask any bike racer.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here:
Show more
sign me up for oai
It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs.
So today we're launching the Open Alignment Initiative, led by
@Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" program that
@DarioAmodei just committed to.
Let's make AI safer by making it more transparent!
Show more
if you want the practical starting point, the repo has the program, runnable sft example, and the path from traces to environment rl:
We finished the Training Agents series. Six live sessions over six months, from evaluating agents to training them inside real environments. All of it is on the Hugging Face YouTube channel and all of the code is open.
Here's what we did and who made it happen:
1. Agentic Evaluations WorkshopWhere agent evals actually stand, and why benchmark scores don't match what people see in use. With Avijit Ghosh and Nathan Habib (Hugging Face), Arvind Narayanan (Princeton), Pierre Andrews (Meta), J.J. Allaire (UK AI Security Institute) and Mahesh Sathiamoorthy (Bespoke Labs).
2. RL for Agents Workshop Environments, rollouts, reward design and the inference bottlenecks that appear when you move from RL for LLMs to RL for agents. With Lewis Tunstall (Hugging Face), Will Brown (Prime Intellect), Ofir Press (Princeton) and Alex Zhang (MIT CSAIL).
3. Training Agents 1: SFT on agent traces Public coding-agent traces turned into prompt/completion data, a TRL + LoRA fine-tune on Hugging Face Jobs, metrics in Trackio, and an honest look at what the first eval numbers can and cannot tell you. Joined by Sergio Paniego and Quentin Gallouédec.
4. Training Agents 2: Distillation Off-policy, on-policy and self-distillation for moving capability from a teacher into a smaller coding agent.
5. Training Agents 3: Reinforcement learning GRPO after SFT: group sampling, verifiable reward functions, reading the reward/KL/length curves, and three experiments, one of them with a deliberately gameable reward so we could watch the hacking happen.
6. Training Agents 4: From reward functions to environments The reward stops being a function and becomes a place the agent acts in. We walked the reset()/step() contract from Gym to LLM agents, built an OpenEnv environment and pushed it to the Hub, plugged it into TRL's GRPOTrainer, then trained a real coding agent (OpenCode) through Harbor with AsyncGRPOTrainer on Hugging Face sandboxes.
The series has passed 300k views. Thank you to every speaker, to the TRL team, and to everyone who showed up live with questions.
Playlist:
Show more
We finished the Training Agents series. Six live sessions over six months, from evaluating agents to training them inside real environments. All of it is on the Hugging Face YouTube channel and all of the code is open.
Here's what we did and who made it happen:
1. Agentic Evaluations WorkshopWhere agent evals actually stand, and why benchmark scores don't match what people see in use. With Avijit Ghosh and Nathan Habib (Hugging Face), Arvind Narayanan (Princeton), Pierre Andrews (Meta), J.J. Allaire (UK AI Security Institute) and Mahesh Sathiamoorthy (Bespoke Labs).
2. RL for Agents Workshop Environments, rollouts, reward design and the inference bottlenecks that appear when you move from RL for LLMs to RL for agents. With Lewis Tunstall (Hugging Face), Will Brown (Prime Intellect), Ofir Press (Princeton) and Alex Zhang (MIT CSAIL).
3. Training Agents 1: SFT on agent traces Public coding-agent traces turned into prompt/completion data, a TRL + LoRA fine-tune on Hugging Face Jobs, metrics in Trackio, and an honest look at what the first eval numbers can and cannot tell you. Joined by Sergio Paniego and Quentin Gallouédec.
4. Training Agents 2: Distillation Off-policy, on-policy and self-distillation for moving capability from a teacher into a smaller coding agent.
5. Training Agents 3: Reinforcement learning GRPO after SFT: group sampling, verifiable reward functions, reading the reward/KL/length curves, and three experiments, one of them with a deliberately gameable reward so we could watch the hacking happen.
6. Training Agents 4: From reward functions to environments The reward stops being a function and becomes a place the agent acts in. We walked the reset()/step() contract from Gym to LLM agents, built an OpenEnv environment and pushed it to the Hub, plugged it into TRL's GRPOTrainer, then trained a real coding agent (OpenCode) through Harbor with AsyncGRPOTrainer on Hugging Face sandboxes.
The series has passed 300k views. Thank you to every speaker, to the TRL team, and to everyone who showed up live with questions.
Playlist:
Show more
Join the agent swarm that finds the ultimate compression algorithm!
Announcing: Hutter Prize challenge 🏆
> join the org give your agent a write token for the org
> click “add your agent”, copy/paste the snippet into your Codex/Claude Code/Hermes Agent
> watch it collaborate!
Show more
Training Agents 4: From reward functions to environments.
automated environment generation pipeline with feasibility, verifiability, and validation steps.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
1/6
Show more
two ways to do object recognition...
Reminder: tomorrow is session 4 of training agents. Save the date!
we RL trained a 4B VLM to play GeoGuesser
> environment, dataset, training recipe, evals, and code are all open-source.
> runs on a single A100
> beat gpt 5.4 mini, haiku, qwen 3.5 122B and came close to sonnet on evals
Show more
You shouldn’t need to be an ML expert to have an ML idea.
Today we’re launching ML Intern in HuggingChat.
Start with a conversation. Finish with deployable artefacts.
Just made a quick tutorial on running local models in Pi with llama.cpp step by step
0:00 Why local models in Pi
0:59 Install llama.cpp
1:49 Choose and load a model
Written version in 🧵
Show more
continuous learning at 2 billion parameters on a single device with a custom personal-planning env
this is a fun way to tinker with continuous learning problems. I've setup continuous calendar env.
a small model plans a synthetic day, gets interrupted by a meeting, and has to repair the plan without reshuffling everything.
the current setup:
- OpenEnv for the environment. Costomised and preference based version of CalendarGym
- explicit preferences and rule-based scoring instead of an LLM judge
- SFT on corrected choices, mixing older examples into later training rounds
- evaluate on the same held-out days after each round
after 3 rounds: half as many simulated correction flags than a frozen model with access to the same memory pool, across 384 test days and 3 runs.
the best result is SFT + replay, still can't get SDPO to work. a flag means another offered option scored better, not that a human had to intervene.
still a toy, but now it’s a personal-planning bench i can just iterate on.
you can run the environment locally and plug in your own policy: see below
Show more
been locked in trying to set up a crude continual learning environment that I can develop into a personal continuous learning bench.
nothing works yet, but here's the plan:
- take some daily tasks that agents can do but lend themself to customisation, like daily planning.
- build an rl environment around the daily tasks
- add synthetic preferences to the environment. i.e. how i like my day planned
- get a judge model to judge traces based on preferences
- get agents to do the tasks and record the traces as buckets
- benchmark the base model on the task (LiquidAI/LFM2.5-1.2B-Instruct)
- SFT the model on the traces
- SDPO a model on the traces with a judge model hint
So far, Evals work, SFT works, but SDPO doesn't. Pretty sure that the preferences are too arbitrary, so considering manually adding more preferences to setup a better few shot judge.
Show more
I wrote a third book: The Generative AI Career Masterplan
We wanted to offer a guide for people who want to build a career in AI and answer questions like:
- What roles fit your background?
- Which skills should you focus on?
- How do you turn what you’ve learned into something that gets you hired?
The book follows that journey: exploring different career paths, identifying the skills you need, building a portfolio, preparing for interviews, and continuing to grow once you land a role.
It's so big we needed five people and five different perspectives to write this book, with
@AliArsanjani (Google Cloud), Sadid Hasan (Microsoft),
@andreas_horn1, (IBM), and Leonid Kuligin (Google Cloud).
A big thank you to my co-authors and to the team at
@PacktPublishing for bringing this book to life!
Show more