登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Lazarz
@Laz4rz
@arcee_ai, before: @EPFL, @ETH, @physics_UW, R6 veteran
参加 September 2017
1.7K フォロー中    7.2K ファン
at this point I don't even know anymore dude you can run DAPO async RL run with vanilla setup or crazy exact batch invariant one that has 2-3x throughput loss and while it fairly crazily reduced the divergence in logprob_diff, it basically doesn't change the validation reward in general a nice read, but offpolicy 0 would be nice just to see baseline behavior also kind of weird they don't see logprob divergence on tmax and r1 runs? also funny cause when you look at the fp16 train-inference mismatch paper, you'll see training absolutely exploding while logprob diff ~0.05, but here its just chillin, some longer nice with similar divergence behavior would be nice to see tho, cause there is no way you can just endlessly let it diverge like this and I dont believe it would find a natural ceiling
もっと見る