People are not looking at the supply side nearly enough.
What matters right now is that making HBM in greater volume and with more superior specs cuts commodity DRAM supply far too much.
This also affects SOCAMM2 supply, which will affect total Rubin rack shipments.
In other words, cutting HBM to increase SOCAMM2 supply and raise total Rubin rack shipments is the rational move, and that fits Nvidia's interest.
The resulting performance loss is solved at the NPO layer and addressed at the cluster level.