註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Perry Dong
@perryadong
CS PhD @StanfordAILab, Part Time @GoogleDeepMind | Reinforcement Learning
加入 August 2015
114 正在關注    1.2K 粉絲
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
顯示更多
0
5
306
29
轉發到社區