注册并分享邀请链接,可获得视频播放与邀请奖励。

Radhika Gaonkar
@GaonkarRadhika
Applied Research @primeintellect | ex-Microsoft
加入 July 2019
1.3K 正在关注    486 粉丝
Excited to be presenting at Ray Summit today! I ran a bunch of experiments digging into reward hacking and strange verifier behaviors that crop up in long-horizon RL training. There’s been a lot of interesting work in this space, and my talk will cover a framework to stress-test these behaviors before you scale your training runs. Come say hi if you're around! All built on @PrimeIntellect’s stack!
显示更多