Sharing my first of hopefully many research blog posts!
This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs.
I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.