註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
加入 May 2023
214 正在關注    153.4K 粉絲
Benchmarking @NVIDIAAI's Nemotron Puzzle 75B locally on the GX10. NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations. 🏃‍♀️22.75 tok/s in a single session into 88.85 cumulative at 7 sessions. 🧍Baseline without MTP: ~16.3. Scripts + setup:
顯示更多