註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Radhika Gaonkar
@GaonkarRadhika
Applied Research @primeintellect | ex-Microsoft
加入 July 2019
1.3K 正在關注    486 粉絲
Excited to be presenting at Ray Summit today! I ran a bunch of experiments digging into reward hacking and strange verifier behaviors that crop up in long-horizon RL training. There’s been a lot of interesting work in this space, and my talk will cover a framework to stress-test these behaviors before you scale your training runs. Come say hi if you're around! All built on @PrimeIntellect’s stack!
顯示更多