Register and share your invite link to earn from video plays and referrals.

Volatile Markets
@volatilemarkts
Primary work in Health Care. Passion for Markets, Technicals, and AI Modeling. Acquire the compute while you can…
1.2K Following    1.6K Followers
I want an M5 Mac Studio!! I run a rack of M3 Ultras, and the FOMO is real. Then @ashhart dropped TensorFold 0.3.0 — THIS MORNING — and crushed it. GLM-5.3-Flash now runs on two DGX Sparks at 1.8–2.1× vLLM — byte-identical output. Tested. Confirmed. The recipe book hands you the kernels, the measured numbers, AND the dead ends. Custom-made for everyone. We all get to pick up what he's laying down. No MLX GLM in TensorFold just yet — so we built one for our Studios. That is how user friendly @ashxhart repos are! Our production M3 Ultra seat now serves GLM-5.3-Flash through TensorFold: 45 → 60 tok/s, first token landing before you blink, output word-for-word exact. Started with @MiaAI_lab recipe and got some help from @Kurcide... It feels like new silicon. It isn't. It's a recipe. Our MLX engine is up as PR #9# — for kind consideration. His repo, his call. We're just the lucky ones running it in production while he looks it over: Don't believe us — get the repo and feel the speed today: Ash is a man of the people — MCDMA, Imprint, TensorFold — out here helping all of us get it. My M5 fund stays in my pocket......maybe...
Show more
IT WORKS! For a year, anyone with DGX Sparks and Mac Studios has lived with the same problem: the NVIDIA boxes are fast at reading, the Apple boxes are fast at answering, and they can't share a thought. Two islands. A 10 GbE cable between them. Earlier this month we published pd-bridge ( prefill/decode disaggregation across NVIDIA and Apple silicon. Spark 1 and Spark 2 run the prefill, a Mac Studio runs the decode, and the KV cache crosses between them — reconstructed on the Mac bit‑exact (313/313 arrays) so the decoder sees a normal prefix‑cache hit and never knows a bridge exists. Up to 3.7× faster than a Mac alone at 241K tokens, over plain 10 GbE HTTP. That was one‑way, one turn, one Mac. Today it's a loop! The question we set out to answer: Can a single conversation circulate across four machines — two DGX Sparks, two Mac Studios — turn after turn, each doing only what it's built for, and beat a single machine doing the whole job? How it works: • Spark 1 + Spark 2 (vLLM TP2) — Prefill. Read the conversation, compute the attention state. Blackwell compute: ~1,100 tok/s vs ~400 on a Mac Studio. • Mac Studio A — Gateway. The only Studio with a Mellanox ConnectX‑4 on the MikroTik fabric (via a Thunderbolt 5 enclosure). Receives the Sparks' cache over RDMA — kernel bypass, no TCP, no files — using @ashxhart MCDMA driver ( and maps it into oMLX cache blocks. First time that driver has run on a ConnectX‑4. First time through a mikrotik 812 switch. 6 µs latency, down from ~400! This is the key and we cannot stress enough how appreciative of his work, efforts, and willingness to give to the open source community! • Mac Studio B (512 GB) — Memory. Holds the entire conversation's cache, receives each new piece from A over Thunderbolt 5 RDMA, and generates the reply. Never re‑reads the history, because it already has it. • MikroTik CRS812 + TB5 mesh — Two kernel‑bypass RDMA fabrics carrying attention state between three machines. The reply becomes part of the next message and goes back to the Sparks. Inference, back and forth, across two kinds of silicon. NVIDIA computes what Apple speaks. The big wins: • Up to 4.2× faster replies than a Mac Studio alone. Alone, Mac Studio B sits at a flat ~31 s per message because it re‑prefills the entire history every time. Fed by the Sparks: .7 s at 31K tokens, 7.4 s at 47K, 10.5 s at 63K. The longer the conversation, the more the split pays. • The Mac never prefills. Cached tokens on Mac Studio B: 30,720 → 47,104 → 61,440. Every block computed on Blackwell, consumed on Metal. • Network transport is a rounding error. Spark → Mac Studio A over RDMA through the MikroTik: 0.3–0.5 s. Mac Studio A → B over TB5: ~0.05 s. Under 1 second of a ~8‑second reply. (Next: routing the cache back to the Sparks so they only prefill what's new — that flattens the whole loop.) Who made this possible: @ashxhart — MCDMA. RDMA verbs on macOS. Nothing here works without it! @Apple @NVIDIAAI @mikrotik_com ( @b_ostrov — MelonDMA, and the KV‑block‑over‑RDMA ideas we built on, and because I love that Ben is using M1-M2 machines and making this backward compatible for anyone/everyone with a card and imagination.. @MiaAI_lab — the model builds we run, because you already know she's going to be the queen of the heterogeneous model scene! Why we're posting: This is a basement, not a lab. We're not selling anything: compute where compute is cheap, memory where memory is cheap: RDMA in between. What if we don't stop here..? Open Source Must Win!
Show more
Running (3) instances of GLM 5.2. (6) Sparks, API, and on up to (5) Mac Studios in a single cluster for the highest fidelity GLM I can reach. GLM can coordinate near persistent memory agents. It’s proto-agi… I have never been more excited about @Zai_org has revolutionized local compute. I cannot wait to see what GLM offers us next. @MiaAI_lab @Prince_Canuma cannot wait for the recipes for my sparks and studios.
Show more
Ox Alpha - No logo - No branding - No PR - Lots of mystery - Lots of rumors A frontier model with 100T free tokens a day. Endless guessing about who is behind it. The anonymity became the distribution and the timeline ate itself alive. This is how they forced everyone in AI to pay attention. IMO the most brilliant marketing move in AI of the year.
Show more
0
154
2.2K
124
Forward to community
Q: Is the high yield credit market jumping up and down screaming "imminent crisis"? A: No
Ahhh.. must be the GLM 5.x Air. Is it going to be GLM 5.2 Air or the GLM 5.3 Air LFG….
Good luck to those short the market. You’re going to need it into the midterms. This remains a melt-up environment. There’s a zero percent chance we’d be short this market. 🚀
This is the first time any access to Mythos 5 has been granted beyond the small group that had access through Project Glasswing.
What IF and I means HUGE IF… If OX ALPHA is a brand new GLM 5.3 “FLASH” model ?
I wanted to see if there was any real difference between the DGX Spark and its siblings. Seven GB10 Sparks, same silicon, different chassis, airflow, fan design, power tuning and sustained behaviour. I compared the HP ZGX, Acer GN100, MSI EdgeXpert, Lenovo PGX, Gigabyte AI TOP ATOM, ASUS GX10 and Dell Pro Max on thermals, sustained performance and noise. • HP ZGX Nano G1n Best overall balance. Peak CPU 77.3°C, GPU 69°C. HP rates it at 27.6 dBA under intensive workloads. My ZGX also runs about 10°C cooler than my PGX under sustained loads. • Acer Veriton GN100 Best thermals. Peak CPU 74.7°C, GPU 69°C, GPU power 69.2W. • ASUS Ascent GX10 Strong thermal design. Peak CPU 87.3°C, GPU 82°C, GPU power 69.8W. • MSI EdgeXpert MS-C931 Performance focused. MSI claims 1,729 tok/s vs ~1,600 tok/s for the reference design, with the rear chassis 15°C cooler, top 9.1°C cooler and SSD 9°C cooler in its own testing over a DGX Spark. • Lenovo ThinkStation PGX Performance focused. Sustained system draw has been measured at roughly 104 to 160W. In longer heat-soaked runs, it can edge out DGX Spark by a few percent. • Gigabyte AI TOP ATOM Aggressive power tuning. Peak CPU 90°C, GPU 81°C, GPU power 75.5W. • Dell Pro Max with GB10 Runs warmer. Peak CPU ~87.7°C, GPU ~80°C, GPU power ~70W. Sustained decode settles around 80 to 83°C CPU and 68 to 71°C GPU. The performance gap between these machines is usually small; the thermal gap is a different story.
Show more
Local AI is how we escape the marginal cost of intelligence.
A year ago the consensus was everyone would be using one or two frontier models, running entirely on NVIDIA GPUs. The reality now is we have hundreds of models each with different cost/speed/intelligence tradeoffs for different tasks, running on dozens of different local / data center chips (even older hardware). Ironically the frontier models that were trained on NVIDIA GPUs have made it faster and easier to port kernels from CUDA to other hardware targets (AMD, Apple Silicon, Intel, ...). There are people with no experience writing kernels using relatively cheap AI agents to implement efficient kernels for Apple Silicon, synthesizing all the tricks from existing CUDA implementations that were meticulously hand written by experts. Funnily enough, since the performance of these kernels can be quickly and objectively evaluated by an agent by actually running the kernel on the hardware, it's an unreasonably tractable task for agents. They can continuously improve them in a fairly simple autoresearch loop. Kernel interoperability is solved ("The unreasonable effectiveness of AI agents"). The next continuation of this trend is to have millions of small, specialized models for each use case, tuned for each user based on how they use the models / their SLAs (e.g. maybe you're fine running a job overnight so it can run on cheaper hardware). They will run efficiently on any hardware, so then it's a choice of which hardware is the best for that specific thing. For this we need better infrastructure to evaluate models and map out the tradeoff space so we can give the optimal point in the tradeoff space of Model x Quant x Harness x Inference Engine x Config x Hardware.
Show more
There is a new HIDDEN “Ox” model showing up on Openrouter and people have said it’s been GREAT. @MiaAI_lab posted this photo and in the comments said there’s a “connection” here. I’m curious if Deepseek is about to drop a 3rd model ? They just dropped a DS4 Flash Vision Today.
Show more
Local AI had a great year & it’s about to get wild!
High probability this is a GLM-5.3 variant with Vision 🔥
I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is 👀
0
750
2.6K
51
Forward to community
The rumours are right. 0xalpha is absurdly good. It figured out an issue Fable and Sol have been kicking back and forth for a few hours now.