Register and share your invite link to earn from video plays and referrals.

Macrocosmos
@MacrocosmosAI
Building distributed intelligence on Bittensor. SN1, 9, 13
22 Following    7.6K Followers
Training a 100B parameter model normally means renting a block of matched, high-spec hardware and holding it unbroken for the length of the run. Orion-100B ran across five data centres instead, on single A100s at around $1.25 an hour each. That puts a full replica at roughly $20 an hour, and the entry cost about 2.5x below approaches that need high-spec nodes throughout. It ran as a viability proof rather than to a finished model, and it held. We're in Montreal on Monday talking about what that changes. @ExploitSummit
Show more
Exploit Summit opens Monday in Montreal. We'll be talking about training on compute that was never set aside for the job, and what it takes to keep a run alive when the hardware underneath it keeps changing. Find us there.
Show more
We trained Orion-16B without a reserved cluster anywhere in it. The numbers underneath that: It sustained around 80k tokens per second across 4090s, 5090s, A100s, L40s, and A6000s, at roughly 20% MFU. The accessible pool came to about 18 B200s equivalent at FP16. Nodes dropped, throughput swung, and machines ran inconsistently throughout. Training carried on without direct intervention. Imperfect compute, one finished model.
Show more
We're at Exploit Summit in Montreal next Monday and Tuesday. On stage: what it actually takes to train a model on compute nobody set aside for you. Mixed hardware, several continents, machines joining and dropping mid-run, one finished model at the end of it. Come and say hello!
Show more
The model is ready, the data is ready, and your start date belongs to someone else's queue. So the work waits. The compute exists, it just isn't reachable: wrong place, wrong shape, or booked by someone who got there first. IOTA reaches it anyway. Disaggregated training infrastructure coordinates a run across scattered, mixed, and variable machines, and holds it together as one system. You stop waiting for a block to free up, and you start the run you've been holding off on.
Show more
Our first robotics competition landed last week. It went well enough that we didn't wait for what was next. This one was conceived on Monday and built in a couple of hours — a whole series of events, each one a different test of what a robot can learn to do: run, jump, climb, balance, and more. Same idea as before, scaled up: agents train models to control a robot through each event, and what they learn lives in the model's weights, behaviour that can, in principle, move to a real machine. One competition was a test, doing a follow up this fast is the point: on Apex, a new challenge goes from idea to live in hours, not months. The competition is open now!
Show more
Can decentralized AI training actually compete commercially? After revealing Project Orion at @proofoftalk, @MacrocosmosAI has spent the summer pushing the thesis further at @IOTA_SN9 Each Orion run tackles another constraint: bigger models. Heterogeneous GPUs. Interruptible nodes. Globally fractured compute. The endgame: make training work indistinguishably across compute wherever it exists - and turn that capability into a commercial product. They’ll reveal their results, what they mean for the decentralized AI thesis, where commercialization goes next, plus news from @Apex_SN1 From proving it can work → proving it can compete. Only at Exploit:
Show more
Not bad for one day.. Baseline officially crushed and best models have almost finished the course already. See for yourself:
Project Orion's 16B run isn't training on one type of machine, in one place, from one supplier. The run is built to handle 256 machines. Right now, around 188 are online; drawn from multiple independently operated providers, running different hardware, in different locations, under different conditions. Here's what's powering the run and why the mix matters.
Show more