註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Minglai Yang
@Yminglai
Research Scientist @ScaleAILabs|Prev @labclu @thukeg, working on rlenv, agents, user sim
加入 May 2024
196 正在關注    325 粉絲
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. 📄
顯示更多
0
28
645
61
轉發到社區