Register and share your invite link to earn from video plays and referrals.

Perry Dong
@perryadong
CS PhD @StanfordAILab, Part Time @GoogleDeepMind | Reinforcement Learning
114 Following    1.2K Followers
World models have emerged as one of the biggest directions in physical AI. At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own Can we get the best of both worlds? We propose Q-Learning with World Models (QWM) (1/7)
Show more
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
Show more