Register and share your invite link to earn from video plays and referrals.

Mia
@MiaAI_lab
Building with AI & LLMs | Insights, recipes, tools & honest experiments
409 Following    35.5K Followers
I want an M5 Mac Studio!! I run a rack of M3 Ultras, and the FOMO is real. Then @ashhart dropped TensorFold 0.3.0 — THIS MORNING — and crushed it. GLM-5.3-Flash now runs on two DGX Sparks at 1.8–2.1× vLLM — byte-identical output. Tested. Confirmed. The recipe book hands you the kernels, the measured numbers, AND the dead ends. Custom-made for everyone. We all get to pick up what he's laying down. No MLX GLM in TensorFold just yet — so we built one for our Studios. That is how user friendly @ashxhart repos are! Our production M3 Ultra seat now serves GLM-5.3-Flash through TensorFold: 45 → 60 tok/s, first token landing before you blink, output word-for-word exact. Started with @MiaAI_lab recipe and got some help from @Kurcide... It feels like new silicon. It isn't. It's a recipe. Our MLX engine is up as PR #9# — for kind consideration. His repo, his call. We're just the lucky ones running it in production while he looks it over: Don't believe us — get the repo and feel the speed today: Ash is a man of the people — MCDMA, Imprint, TensorFold — out here helping all of us get it. My M5 fund stays in my pocket......maybe...
Show more
TensorFold is going to change everything 🚀 Exciting times ahead! ⏳
TensorFold Inference Engine is here 🚀 I spent six months making one weight read count for more than one token on Apple Silicon. Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes. Qwen 3.8 27B MLX 4Bit - 120-124tks Nemotron Lightning MLX 4Bit - 188-206tks Qwen3.8 Flash Next MLX 4Bit - 88-92tks CUDA Implementation is in Alpha showing strong gains. The Repo is in the comments 👇🏼
Show more
Some stats creating this with Opus 5.5 set on Ultracode: Most of the time there were between 30-45 agents running at the same time. Total cost: $415 Total weekly limit used on 20x max plan: 25% Total tokens burnt: 1.2 billion 🤯
Show more
Built by Opus 5.5 (Ultracode) Created by its own creativity. No images, no video, no audio files. Every frame and every note is generated by code, live, from the word you type. Live here: How it works and the full prompt below.
Show more
Qwen3.8 Flash for 2x DGX Sparks just got better 🔥 - Faster decode on prose and code. - Faster follow-up replies. - Optional vLLM 0.30 path. Get it here:
@Alibaba_Qwen Qwen3.8-Flash-Next on two @NVIDIAAI DGX Sparks just got better 🚀 It inherited some of the improvements i did on the single-spark recipe and a little extra more! Default recipe 👇 ・Multi-turn fix: follow-up replies 5x faster (3.25s → 0.63s), after tool calls 2.55s → 1.98s ・Code up to +27% faster (333 → 422 tok/s at 8 streams) ・~56 tok/s prose, ~76 code (1 stream), up to ~226 prose / ~422 code (8 streams) ・Optional bit-exact decoding across both nodes ・Cached-token reporting in the API ・Safer memory default: no more running both nodes at under 1 GiB free New opt-in lane on vLLM 0.30 (./start-v030.sh) 🌟 ・Follow-ups in 0.28-0.40s, and repeated prompts hit the cache (0.35s vs 2.7s) ・Same decode speed, steadier: fixed a vLLM 0.30 default that made speed swing ±15% between runs ・One small patch instead of seven ・Optional FP8 KV: 2.6M tokens of context, 3/3 needles at 200k Tested for stability for a few hours! I will keep testing this over the next days and probably make the vLLM 0.30 script the default one in the next version of the recipe ✌️
Show more
Built by Opus 5.5 (Ultracode) Created by its own creativity. No images, no video, no audio files. Every frame and every note is generated by code, live, from the word you type. Live here: How it works and the full prompt below.
Show more
I see a lot of finger-pointing in the AI community, and it probably doesn’t help anyone. My mantra: A rising tide lifts all ships. @MiaAI_lab has credited me several times for tool-eval-bench even though she could have just pointed to a fork or her own project. Instead, she contributed to my project to help make it better for everyone. Let’s not forget that we’re all trying to make things a little easier and better for everyone.
Show more
The numbers are in. Most own 2x DGX Sparks, followed by a single spark. Only 7.2% own exactly three. More than interesting: over 25% own at least 4x DGX Sparks Total 1,796 votes Total "sparks" participated: at least 4,137 A lot of people commented "where is the 0 option".
Show more
How many DGX Sparks do you own?
Opus 5.5 Ultracode still running. Now it's at the "let-there-be-polish" phase, a complete different sub-task running 21 agents. Will post results when it's done!
Opus 5.5 with Ultracode (had to restart) Currently running 32 agents, with 9.2M tokens burnt.
Fully support @MiaAI_lab She has done incredible work for dgx open source community These attacks make no sense at all !! Keep doing great work !!!!
A living legend owns 36 DGX Sparks 🔥🤯
I’m building “The All Spark” a 36x DGX Spark cluster to run my own local agents and support the community with compute. 24 Sparks running today on a single cluster with the rest coming online after a quick power upgrade to the house 😅
Show more
Opus 5.5 with Ultracode (had to restart) Currently running 32 agents, with 9.2M tokens burnt.
I feel sorry for UK. They published "rules" on how to use AI: Basically saying to use AI only when needed, with short prompts and as few interactions as possible to reduce environmental impact 🤣 Will probably be adopted by EU too.
Show more
If people think this is a joke - it’s not:
Useless in what? For being completely private? For having no limits? For not being nerfed after a while? For having no downtime? For being able to own intelligence instead of renting one? Yes, without all of these local models are completely useless I agree 💯
Show more
After the new additions to the Qwen3.8-Flash-Next for a single spark, coming with similar updates for the dual recipe! Testing rn to ensure it's stable. Dropping possibly tomorrow!
Show more
How many DGX Sparks do you own?
Come check out my website. It includes a list of my latest & greatest recipes. The "Install Model" button gives you a ready-to-use prompt for Claude, Codex, or others. It adapts the recipe to your setup, starts the server and tests it.
Show more
Claim $250 in FREE cloud credits for Claude. Link below 👇
Whether you like it or not, that's the truth... Opus 5.5 is just wow
let’s be honest Anthropic really cooked here…
The real test for Opus 5.5 is what happens a week or two from now. Will Anthropic follow their own tradition of nerfing models or will they surprise us?
btw I've been trying so hard to reach 5h limits on my 20x max sub, but no matter what I do, I can't! Strange thing to to say, but limits are actually generous now, including the weekly.
Oh this is nice... got a reset button in Claude Abusing my limits now