註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Chelsea Finn
@chelseabfinn
Asst Prof of CS & EE @Stanford Co-founder of Physical Intelligence @physical_int PhD from @Berkeley_EECS, EECS BS from @MIT
加入 June 2014
397 正在關注    102.5K 粉絲
Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch. We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining. Paper:
顯示更多
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
顯示更多
0
13
645
54
轉發到社區