登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Perry Dong
@perryadong
CS PhD @StanfordAILab, Part Time @GoogleDeepMind | Reinforcement Learning
参加 August 2015
114 フォロー中    1.2K ファン
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
もっと見る