Training and rollout logprobs matched bit for bit on ROCm. The
@RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime. A 200-step Qwen3-8B GRPO run on 8×
@AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout.
The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides.
Deep dive: