I want an M5 Mac Studio!!
I run a rack of M3 Ultras, and the FOMO is real.
Then
@ashhart dropped TensorFold 0.3.0 — THIS MORNING — and crushed it.
GLM-5.3-Flash now runs on two DGX Sparks at 1.8–2.1× vLLM — byte-identical output. Tested. Confirmed. The recipe book hands you the kernels, the measured numbers, AND the dead ends. Custom-made
for everyone. We all get to pick up what he's laying down.
No MLX GLM in TensorFold just yet — so we built one for our Studios. That is how user friendly
@ashxhart repos are! Our production M3 Ultra seat now serves GLM-5.3-Flash through TensorFold: 45 → 60 tok/s, first token landing before you blink, output word-for-word exact. Started with
@MiaAI_lab recipe and got some help from
@Kurcide...
It feels like new silicon. It isn't. It's a recipe.
Our MLX engine is up as PR #
9# — for kind consideration. His repo, his call. We're just the lucky ones running it in production while he looks it over:
Don't believe us — get the repo and feel the speed today:
Ash is a man of the people — MCDMA, Imprint, TensorFold — out here helping all of us get it. My M5 fund stays in my pocket......maybe...