注册并分享邀请链接,可获得视频播放与邀请奖励。

filipe
@filicroval
data eng | serendipity business inquiries: filipe@filicroval.com
加入 May 2023
297 正在关注    153.5K 粉丝
i added native support for @NVIDIAAI's Nemotron Puzzle 75B to mlx-lm. it now runs natively on an M2 Max 64GB: ⚡️22 tok/s 💾45.5 GB peak memory usage 📚4-bit experts + 6-bit dense + BF16 head i also fixed an annoying numerical bug in mlx-lm. outputs were subtly wrong, cosine similarity was 0.8832 vs NVIDIA's reference (identical inputs). the culprit was one dtype cast in the Mamba layers happening in a different spot than NVIDIA's. once moved, the cosine similarity improved to a satisfying level (0.999...). related PR: weights:
显示更多