註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Deli Chen
@victor207755822
Deep Learning Researcher @deepseek_ai | #AGIforEveryone# Prev. BS and MS @PKU1898 | | All opinions are my own. | INTP-T | 人心惟危,道心惟微
加入 December 2023
181 正在關注    31.7K 粉絲
Just my opinion: The real reason PPO can handle long-horizon tasks? The Value Model for multi-step task. Honestly, it just shifts the training difficulty from GRPO’s GRM onto the Value Model. I mean, the problem isn’t solved — it’s just been relocated. The core question remains: How do you get stable process supervision in long-horizon tasks? That’s the real bottleneck.
顯示更多
0
0
165
10
轉發到社區