Currently the M5 Ultra is very much a competitor but don't count the Sparks out. They hold their own if serving more than one user.
Tonights testing got some very nice single user speeds using Tensorflow, but as expected prefill is important.
I ran 3 tests on long context windows.
1. M5 ultra+ Tensorflow
2. M5 Ultra plus oMLX.
3. 4x DGX Spark vLLM.
All on GLM 5.3 Flash.
The Sparks blow the M5 out of the water on cold prefill.
When in a coding session and the context window grows over time large context windows are fine on the M5, even with oMLX.
I'm still tracking down a serious bug with Tensorflow slowing down considerably after a long context window fills the ram. So these results may change after that gets squashed.
More to come.
顯示更多
First "Big" finding on the M5 Ultra.
TLDR: GLM 5.3 Code is running at 125 tok/s and prose at 87!
Earlier in the day, I optimized oMLX for GLM5.3 and peaked at a very respectable 57 tok/s prose which is competitive with my 4x Spark cluster.
After running Tensorflow and some tweaking for M5, (comments added to the PR for GLM5.3 so you can reproduce) I'm up to 87 tok/s prose and 125+ for code.
Now, this is not really fair as I haven't ported Tensorflow over to the 4x Spark stack yet, so I don't know what gains I'll see there. That's coming next.
Being able to load / unload / test / try so darn quick on the M5 Ultra turned what would have been a full days work into just a few hours thanks to the incredible speed of the M5 Ultra's RAM.
More to come but right now, the $10k for the Mac Studio is money well spent.
顯示更多