I pulled the recipe file behind this. The gap has a name, and it is four lines long.
vLLM's official DeepSeek V4.1 Flash recipe landed 10 Sept at 06:32 UTC. AMD was in the hardware block on that very first commit, mi350x verified next to h200, gb200 and gb300, with its own image on the same 0909 build tag. That is the vLLM image, a separate artefact from whatever AMD ships on its own hub, so both things can be true at once.
18 hours later a commit signed off from an AMD email address added one environment variable, plus a comment saying why:
DeepseekV41ForCausalLM does not support torch.compile. The ROCm sparse SWA backend only reports UNIFORM_BATCH, so default FULL_AND_PIECEWISE cannot start unless breakable CUDA graphs are on.
So the AMD path serves this model with FULL_AND_PIECEWISE off. vLLM's own docs call that mode the default and generally the most performant setting, especially for low latency with small models or MoEs. The flag standing in for it is marked Experimental in and defaults to off. V4.1 Flash is an MoE running 8 to 16B active, benchmarked down at the low-interactivity end. That is the exact corner where the setting bites hardest.
None of which shrinks the 14.8x. It dates it. The V4 generation went from every Instinct SKU marked unsupported, blocked on a Hopper-only kernel launch primitive baked into the model's own TileLang path, to all three verified. Same file, same repo. So the open question is how many weeks this one holds.
Now the part I actually trade. I pulled all 190 recipes in that repo. 146 carry a hardware block, 72 list an Instinct part, and 70 of those are verified. Where one file names both vendors it is 61 of 63. At 100B parameters and up, 46 of 48. Whether a frontier model runs on AMD is close to settled. How fast it runs is not, and that is where the entire fight now sits.
The memory bill sits outside that fight. On this recipe's own verified list, HBM per GPU runs H200 141GB, GB200 186, GB300 278, MI350X 288. The part losing the tokens-per-dollar argument is carrying the most memory per die. DRAM gets paid per gigabyte installed, not per token delivered, and it does not care which logo wins the kernel race. $MU $SNDK
$NVDA $AMD
Show more
POWER OF CUDA MOAT ALERT🚨: 2 days after CUDA vLLM supported DeepSeekv4.1 Flash, AMD finally publicly released its DeepSeek v4.1 Flash image. Functionally, it works out of the box, but performance-wise, it is currently up to 14.8x worse perf per dollar than H200 and up to 42x worse perf per dollar than B200/B300 currently.
The 🚀 POWER OF THE CUDA MOAT 🚀 is that NVIDIA's collaboration with its massive 6 million-developer community ecosystem means that CUDA is optimized on day 0. As AMD Anush said, "Speed is the Moat," and day 0 model support shows CUDA is the speed.
Show more