Register and share your invite link to earn from video plays and referrals.

Trapit Bansal
@TrapitBansal
AI Research @Meta, Founding Member TBD Lab | Previously @OpenAI, co-creator of OpenAI o-series models (thinking in LLMs)
Joined January 2010
283 Following    31.8K Followers
Fun fact: RLSlow was named after Thinking, Fast and Slow. The idea was that language models already had a kind of "fast" thinking, producing an answer immediately, and that we could use RL to teach them "slow" thinking: deliberate, multi-token reasoning that spends more compute working through a problem. What I remember most from those days is how early the team developed real conviction in the direction, and how much work went into earning it. We were developing the algorithms, designing careful experiments to test the ideas, and watching the empirical evidence accumulate. These were many long nights babysitting runs, understanding what the results were telling us, and figuring out what to try next. A lot of those early discussions were with Ilya, and later with Jakub. Pretty early, the evidence had already pushed us to a strong view: RL for reasoning would scale. Models would learn to spend more compute at inference time to reason through increasingly hard problems, and this would fundamentally change how we think about inference. Many of us also spent countless hours reading through reasoning traces. Seeing how models arrived at answers, not just the answers themselves, felt like a powerful new lens on generalization and alignment. These ideas feel obvious now. They really weren’t then.
Show more
Let me take this opportunity to pay tribute to the RLSlow team :) It was a priviledge to lead it through the years, first with @ilyasut then with @merettm and finally on my own - the best team you can ever wish for. It feels like ages ago when we started working on the foundations of reasoning (later berry) with @TrapitBansal, Francis Song, @_smileyball and @sidorszymon. Many more things happened since then, huge thank you to all teammates through the years: Peter Zhokhov, @unsafePerformIO @vladi_petroff @masonmeyer_ @afrozenator @lukasz_kondr @dmdohan @wang12tao @TheRealRPuri @edwardjhu @eqhylxx @suchirbalaji and Dan Selsam for all the help. I hope this great technical work continues to be truly and broadly beneficial.
Show more