๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Shrinivas Ramasubramanian
@stablegradients
๊ฐ€์ž… September 2021
1.8K ํŒ”๋กœ์ž‰ ์ค‘    404 ํŒฌ
Is RL optimizing the right objective? ๐Ÿค” Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies ๐Ÿฅฒ We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. ๐Ÿงต 1/n
๋” ๋ณด๊ธฐ