登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Chinmay
@ChinmayKak
22. gradient ascender. prev RL @MSFTResearch . love @teamIvLabs. dms open!
参加 July 2021
1.7K フォロー中    4.3K ファン
This is quite an elegant way to increase the probability of low-probability, high-reward trajectories. TailRL changes the credit assignment so that trajectories reaching reward levels that few other rollouts reach get more weight. Instead of only pulling up expected reward, it optimizes the probability of exceeding reward thresholds across the whole distribution, with a clean finite version of the gradient estimator and a nice best@k interpretation. It also reduces exactly to MaxRL when rewards are binary. The paper is very well written and has a ton of experiments across different domains. larger rollout budgets during training incorporate progressively higher order best@k , while the tail-focused objective increases the probability of sampling low prob but correct trajectories that become especially useful when you can sample many rollouts at inference.
もっと見る