legendary drop from NVIDIA: ModelOpt 0.45.0
biggest additions:
- New NVFP4 (W4A16) weight-only quantization format that requires no calibration
- Better MoE support (including mixed NVFP4 + FP8 recipes for models like Nemotron)
- Easy MXFP4 → NVFP4 conversion for models like DeepSeek V4 and GPT-OSS
- Various improvements for large-scale PTQ and Megatron workflows
looks like NVIDIA is pushing harder on making 4-bit inference more practical and calibration-free.