๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
๊ฐ€์ž… May 2023
214 ํŒ”๋กœ์ž‰ ์ค‘    153.4K ํŒฌ
i added native support for @NVIDIAAI's Nemotron Puzzle 75B to mlx-lm. it now runs natively on an M2 Max 64GB: โšก๏ธ22 tok/s ๐Ÿ’พ45.5 GB peak memory usage ๐Ÿ“š4-bit experts + 6-bit dense + BF16 head i also fixed an annoying numerical bug in mlx-lm. outputs were subtly wrong, cosine similarity was 0.8832 vs NVIDIA's reference (identical inputs). the culprit was one dtype cast in the Mamba layers happening in a different spot than NVIDIA's. once moved, the cosine similarity improved to a satisfying level (0.999...). related PR: weights:
๋” ๋ณด๊ธฐ