We’re thinking about building more end-to-end RL environment recipes like this 👀
What would you guys want to see next?
Any interesting problems / environments we should try?
You can literally RL a 4B VLM to ace GeoGuesser 🌍
> Open-source code, RL environment, dataset, training setup, evals, and everything you need to reproduce it end-to-end
> dropping soon!!