LLMs are good at CUDA because the internet is full of it. But a model that gives you highly optimized CUDA may still struggle to write compilable HIP.
We built a synthetic data pipeline with multi-agent search and post-trained a 14B open-source model with SFT + GRPO RL, leading to substantially better HIP compilation + correctness rates on AMD MI350X GPUs.