注册并分享邀请链接,可获得视频播放与邀请奖励。

Paolo Rosson
@redp314
Head of Applied AI at Dext. Building | Quantum physics PhD from Oxford | Former Italian blindfolded Rubik’s cube record holder
加入 May 2015
2.3K 正在关注    1.7K 粉丝
Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s. Tuned llama.cpp: ~21. Raw mlx-vlm: ~10, not optimized yet. Numbers below if you're setting it up 👇
显示更多