注册并分享邀请链接,可获得视频播放与邀请奖励。

Minglai Yang
@Yminglai
Research Scientist @ScaleAILabs|Prev @labclu @thukeg, working on rlenv, agents, user sim
加入 May 2024
196 正在关注    325 粉丝
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. 📄
显示更多
0
28
645
61
转发到社区