Benchmarking
@NVIDIAAI's Nemotron Puzzle 75B locally on the GX10.
NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations.
🏃♀️22.75 tok/s in a single session into 88.85 cumulative at 7 sessions.
🧍Baseline without MTP: ~16.3.
Scripts + setup: