Excited to be presenting at Ray Summit today!
I ran a bunch of experiments digging into reward hacking and strange verifier behaviors that crop up in long-horizon RL training. There’s been a lot of interesting work in this space, and my talk will cover a framework to stress-test these behaviors before you scale your training runs. Come say hi if you're around!
All built on
@PrimeIntellect’s stack!