at this point I don't even know anymore dude
you can run DAPO async RL run with vanilla setup or crazy exact batch invariant one that has 2-3x throughput loss and while it fairly crazily reduced the divergence in logprob_diff, it basically doesn't change the validation reward
in general a nice read, but offpolicy 0 would be nice just to see baseline behavior
also kind of weird they don't see logprob divergence on tmax and r1 runs?
also funny cause when you look at the fp16 train-inference mismatch paper, you'll see training absolutely exploding while logprob diff ~0.05, but here its just chillin, some longer nice with similar divergence behavior would be nice to see tho, cause there is no way you can just endlessly let it diverge like this and I dont believe it would find a natural ceiling