First "Big" finding on the M5 Ultra.
TLDR: GLM 5.3 Code is running at 125 tok/s and prose at 87!
Earlier in the day, I optimized oMLX for GLM5.3 and peaked at a very respectable 57 tok/s prose which is competitive with my 4x Spark cluster.
After running Tensorflow and some tweaking for M5, (comments added to the PR for GLM5.3 so you can reproduce) I'm up to 87 tok/s prose and 125+ for code.
Now, this is not really fair as I haven't ported Tensorflow over to the 4x Spark stack yet, so I don't know what gains I'll see there. That's coming next.
Being able to load / unload / test / try so darn quick on the M5 Ultra turned what would have been a full days work into just a few hours thanks to the incredible speed of the M5 Ultra's RAM.
More to come but right now, the $10k for the Mac Studio is money well spent.