Register and share your invite link to earn from video plays and referrals.

Minglai Yang
@Yminglai
Research Scientist @ScaleAILabs|Prev @labclu @thukeg, working on rlenv, agents, user sim
196 Following    325 Followers
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. 📄
Show more