๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
๊ฐ€์ž… May 2023
214 ํŒ”๋กœ์ž‰ ์ค‘    153.4K ํŒฌ
Benchmarking @NVIDIAAI's Nemotron Puzzle 75B locally on the GX10. NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations. ๐Ÿƒโ€โ™€๏ธ22.75 tok/s in a single session into 88.85 cumulative at 7 sessions. ๐ŸงBaseline without MTP: ~16.3. Scripts + setup:
๋” ๋ณด๊ธฐ