가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
가입 November 2016
490 팔로잉 중    1.4K 팬
Been thinking about this for a while, as tasks go to more and more turns and longer horizons, PPO is much more elegant for giving dense, per-turn reward. And here we go.
We’re going back to PPO everybody. Curious if there’s some research or experiments comparing the stability with GRPO style setups when you have compacted rollouts