๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Minglai Yang
@Yminglai
Research Scientist @ScaleAILabs๏ฝœPrev @labclu @thukeg, working on rlenv, agents, user sim
๊ฐ€์ž… May 2024
196 ํŒ”๋กœ์ž‰ ์ค‘    325 ํŒฌ
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. ๐Ÿ“„
๋” ๋ณด๊ธฐ