Register and share your invite link to earn from video plays and referrals.

dylan ツ
@demian_ai
growth @nebiustf @nebiusai // ex @Scaleway // from silicon to token, inference and anything in between. Views are my own - not financial advice
2.6K Following    29.4K Followers
Coming from the goat, means a lot 🐐
Every Muse user is supposed to get their own cloud PC. 2 vCPUs, 8GB RAM, 100GB disk. If Meta actually leaves those boxes on, user growth turns into a chip and memory problem. that’s what i'm trying to size a few thoughts on $META and Muse: Muse is 13 days old and US only App Store downloads are still under 1M on iOS. Add WhatsApp, web, android, mac and you might have 1-2.5M accounts. Most of those are not hot machines. Call live VMs closer to a million Ok, where does this go: - stays messy and US-heavy: maybe 30M in a year - whatsApp works: 25M in six months, 100M in a year - they force it through whatsapp and insta: 250M Let's go with 100M, the middle scenario Those EPYCs have 126 cores. Two vCPUs per user means about 60 VMs per chip if nobody is sharing. Share them and you fit more, and it is less of a private machine. if you leave them 1:1 and always on: today → tens of thousands of chips, 14 PB RAM 25 million → 400k chips, 200 PB 100 million → 1.6 million chips, 800 PB RAM, ~1.6 GW 250 million → 4 million chips, 2 EB 1 billion → 16 million chips, 8 EB $AMD ships on the order of 10 million server chips this year. The whole industry about 39 million. 100 million dedicated Muse users is a visible piece of AMD’s year, and they were already getting called sold out. Next die has more cores, which is how you stuff more VMs on the same chip. CPU time can be shared. That RAM stays reserved if the files are supposed to live on the box. 2026 DRAM is already spoken for, HBM first. 100M users would lock up about 1% of a year’s DRAM bits. A billion users, 10%. Meta’s own power plan is 7 GW this year, 14 next. Training and ads are already in that. The paid tiers do not cover it. Ten million people on $20 is $2.4 billion a year. Last quarter Meta did $60.8 billion, almost all ads, and $31 billion of capex. Full year capex is $130-145 billion. 100 million VMs is $30-50 billion of hardware if you actually build them. So either a lot of those VMs get frozen when idle, or this does not scale as advertised
Show more
Damn it’s crazy how we went from being totally unknown to being mentioned in that list in less than a year Good time to go give @nebiustf a follow
Prediction: within 12 months, top three models will be open source. Economic winners will be the American clouds that serve them: Nebius Iren Baseten Together Fireworks
Is Stripe the biggest product trap of the 21st century? Let’s say you need to get paid, or you just made $100. Stripe is already there. Docs, checkout, webhooks. Naturally you flip it on. The page says 2.9% + 30¢. Sounds fine. So that’s the $100. Then the receipt starts talking: - 2.9% + 30¢ → $96.80 - international card +1.5% → $95.30 - USD in, EUR out, conversion ~1% → $94.30 - Adaptive pricing: customer sees a local price with a 2–4% markup, you still pay processing - Managed Payments +3.5% → $90.80 - Billing extra if it’s a sub, 0.7% → $90.10 - Tax: 0.5% → $89.60 You thought you made $100. You’re holding about $90 and that’s the good version, nobody disputed yet. Then the dashboard isn’t a payments screen. it’s a (low quality) store: - Radar Lite “included” (wtf even is this) - Radar Standard $10–20/mo, or cents per screened payment, for the real AI fraud thing - Smart Disputes auto-fights and takes 30% if it wins - you want one number the reports don’t have → Sigma, $15/mo - Data Pipeline $65/mo to dump it into Snowflake - Signals, Identity, Atlas, Climate, Capital, Issuing, Treasury I opened it to check a payout and spent ten minutes in tabs I never turned on. Then the customer disputes. “Product unacceptable.” You delivered, but it doesn’t matter. - bank files - they take the $100 back - $15 / €20 dispute fee the same day, win or lose - you already paid $10 in processing. they keep that - fight and lose → another $15 / €20 Same $100 sale, worst path: $100 − $10 in stacked fees − $100 clawback − $15 dispute − $15 for fighting What’s left of the $100: about −$40 Negative balance, if that payment was most of what you had. Small ticket or big ticket, it's the same mechanic. Bank can take three months. A friend told me this years ago. I only got it on the ledger: don’t ever run a large payment through Stripe. Why you still use it: - PayPal does the same thing, uglier, and freezes you - Adyen is cheaper at volume and takes weeks of sales - Square is a till - Mollie is fine if you’re small and EU - Paddle / Polar take ~5% to wear the tax - Lemon Squeezy is Stripe now - Visa/MC wrote the chargeback rules for all of them So you stay. You bake 8-12% into the price. You treat a card like it might come back as a debt. Big invoices go on a transfer. A market this size should have five serious options you can switch to asap
Show more
> be intel >spend ten years getting clowned cuz the rack is all GPU > xeon is the chip that boots the box and then sits there > agents drop > not one prompt. 20 tool loops > instinct hits 100k users > muse ships > muse: 1 linux vm per user > browser + shell + cron + sub-agents > free = 100m tokens / week > $20 = 500m > $100 = 3b > GPU still runs the model > CPU still runs the age > GPU doing overtime > still needs a babysitter > babysitter = cpu > every lab that bought blackwells now wants a pile of hosts too > Tan on a splunk stage like > "yeah we can fill about half of you lmao" > CEO groupchat going “send boxes” > memory also 5-7x in the same breath > Q2 server already brrr > xeon 6 actually shipping > timeline still posting GPUs > CPUs joined the chat anyway > AMD already in the replies > no dimms / no watts = no rack > so either agents stay on and Intel prints hosts for 2 years > or this was just rationing copypasta $INTC
Show more
$INTC Lip-Bu Tan admits Intel can fill only half its CPU orders and apologizes when CEOs call him "Yeah, I think inferencing is very important. When you want to drive some of this reinforced learning, and also in terms of agentic AI, CPU is the best." "GPU is very good for training, but CPU is really good for orchestration, control plane, and driving some of this even single thread rather than multi-thread has become very useful." "It happened that I invest quite a few companies in the frontier model. They all tell me that, 'Lip-Bu, we need more CPU.'" "So good news is, right now, sometimes in life you need some help. And so the help that come to me is that CPU is so high demand, I only can provide 50% of what the customer want. So many CEO call me up, I had to apologize, I'm not manufacturing enough for them."
Show more
@GavinSBaker Love to see it, long live open source
Name a better hackathon venue…i’ll wait @hackbarna
Typesafe will be a $100b company in 10 years
Instinct will be a $ 100b company in 10 years.
we live in the dumbest timeline, maybe we do deserve to all die at the hands of intelligence
$20k grand prize is insane
Show us what you’re building 👀 Join the NVIDIA × Nebius Global AI Hackathon powering agents, ai apps, coding tools, physical ai... $50K+ in prizes. Bring your best projects. 42 days left! link below
Show more
The labs committing to gigawatts are still looking for capacity in tens of megawatts. That is a more interesting signal for Nebius than another announcement about the eventual size of the AI market. CNBC reports that Anthropic has explored 20–30 MW deployments in the UK and Nordics, while OpenAI has explored smaller deployments in the Nordics. My read: customers are buying two things when they buy compute. Processing capacity, and the ability to start using it at a particular time. Inference makes that second dimension especially valuable. A large training run needs many GPUs communicating closely. An inference fleet can serve separate requests through independent copies of a model. Each copy may still require substantial, tightly connected infrastructure, but separate copies can operate in different locations. That changes which sites are economically useful. A pocket of available power that cannot support a giant training cluster may still support meaningful production inference. An operator with several suitable locations has more opportunities to match a customer’s workload and deadline. This is the strongest argument for the distributed part of Nebius’s strategy. Its four announced UK deployments are expected to reach 65 MW combined in 2027. They sit alongside much larger projects, including the planned 1.2 GW Pennsylvania campus. Together, these give Nebius several ways to add capacity as demand develops. There is already a commercial signal behind this. For a customer constrained by compute, waiting can mean delayed launches, tighter usage limits and demand it cannot serve. A lower future infrastructure price has to be weighed against those costs. The advantage still has to be earned. A 20–30 MW allocation can sit inside a larger campus, so this reporting does not establish that customers prefer smaller buildings. And distributing capacity creates its own problems: duplicated model copies, uneven demand, more operational work and limits on where customer data can go. Nebius has to keep those sites well utilized and deliver consistent performance. Otherwise, a broader footprint simply becomes a more complicated one. But the strategic logic is strong: regional deployments create additional opportunities to serve demand while larger projects progress. For us, the opportunity is to make more of the world’s available power useful to customers sooner. For those customers, the cost of compute includes the cost of waiting.
Show more
almost every high performance motor on earth turns on a sintered rare earth magnet: airpods Phone haptics EV drive units robot wrists missile fins cmpressor fans in a hall full of GPUs the bottleneck in three numbers: - China mines about 60% of the magnet rare earths - it refines about 91% - it makes about 94% of the sintered permanent magnets having deposits might be a comforting idea, it's still just a small part of the strategy. Ore is not the product. The product is basically a brick that keeps its field when the motor gets hot. Building that brick means separating periodic table twins, alloying them, powdering, sintering, machining, magnetising. China built that middle when the west treated it like a dirty job nobody wanted at home. so who owns separation, and a sinter line a customer will actually trust? A few names sit on those steps: - $MP : Fort Worth metal + magnets, 10X campus behind it - $LYC (Lynas) : the non-Chinese oxides that already ship Energy Fuels + VAC: magnet craft bolted onto a US feed - Noveon: already shipping sintered magnets from Texas - Japan (Proterial, Shin-Etsu, TDK): the houses that never forgot how Educational, not investment advice
Show more
almost every high performance motor on earth turns on a sintered rare earth magnet: airpods Phone haptics EV drive units robot wrists missile fins cmpressor fans in a hall full of GPUs the bottleneck in three numbers: - China mines about 60% of the magnet rare earths - it refines about 91% - it makes about 94% of the sintered permanent magnets having deposits might be a comforting idea, it's still just a small part of the strategy. Ore is not the product. The product is basically a brick that keeps its field when the motor gets hot. Building that brick means separating periodic table twins, alloying them, powdering, sintering, machining, magnetising. China built that middle when the west treated it like a dirty job nobody wanted at home. so who owns separation, and a sinter line a customer will actually trust? A few names sit on those steps: - $MP : Fort Worth metal + magnets, 10X campus behind it - $LYC (Lynas) : the non-Chinese oxides that already ship Energy Fuels + VAC: magnet craft bolted onto a US feed - Noveon: already shipping sintered magnets from Texas - Japan (Proterial, Shin-Etsu, TDK): the houses that never forgot how Educational, not investment advice
Show more
the best model to come out this week is one that cannot write meet Jev they had a nice Wikiracing demo in the release paper. it's a game where you start on one Wikipedia page and try to reach another by only clicking links that already exist on the page. each hop: hundreds to thousands of Wikipedia links. pick one. do not invent a URL. repeat. - LLMs write a link and sometimes hallucinate a dead end (the error compounds) - Jev takes the current state, looks at the options you already defined, and returns a typed choice with probabilities. and the results speak for themselves zoom out and you can feel why this might be a paradigm shift in how we use models. A lot of software needs a reliable fork: which tool, which API, which row, which next action. If that fork gets cheap, typed, and fast, you do not run one careful agent, you spray decisions into the hot path. same shape as Wikiracing, applied to tool routing and guardrails
Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Show more
okay enough stonk talk, time for GPU talk $NBIS we just ran our widest MLPerf Inference round yet and the full rack Nebius system with NVIDIA GB300 NVL72 took first on DeepSeek R1 in both server and offline (about 603k and 690k tokens a second). simple version of what happened: 1. we got the new chips (including Vera Rubin, the latest NVIDIA iron, and we were one of only two labs to show it in this round) 2. we racked them (full GB300 NVL72, plus the smaller boxes people actually buy). 3. we put them to the test under MLPerf, the public inference benchmark where somebody else holds the stopwatch result: the full rack took #1# on DeepSeek R1 for live and batch traffic, and pushed gpt-oss 120B past a million tokens a second. When we grew from 8 GPUs to 72 (almost 9x!), speed scaled almost 9x too. One fast GPU is easy, a rack that stays fast when you multiply it is the product. another day another milestone
Show more
Five first-place results in MLPerf® Inference v6.1. 603,023 tokens/sec on DeepSeek R1 at 72 GPUs and preview-category results on the Nebius @nvidia Vera Rubin NVL72. We were one of only two submitters with results on that hardware. Full results: #MLPerf# #MLCommons# #Inference#
Show more
Nebius just joined @joinstationf F/ai program as a partner for the next cohort Joining alongside great names like Eleven Labs, Openrouter, Github, Hubspot, and Rippling. Station F is the biggest startup campus in Paris 🇫🇷 F/ai is the AI only track inside it: reco only, no open application, built to push early technical teams toward real revenue. the room already had OpenAI, Anthropic, Mistral, Google, Meta, Microsoft, AWS, plus Tier-1 capital like Sequoia and General Catalyst, so a european founder is not cold emailing their way into each door one by one. Cohort one ran that setup, cohort two is where @nebiusai walks in 💪 what is happening is pretty concrete. we already run GPU capacity in Paris. F/ai puts that capacity in the founder layer of the french and wider european AI scene, next to the model and capital partners who were already there. The program is also opening beyond pure Paris residency to hybrid teams across european hubs, and adding a week in SF with partner HQs. it's important bc Europe keeps producing strong research and still loses time (and sometimes companies) on the dull middle of the stack, quotas, infra help, distribution. A program like this with models, compute, tools, and capital in one coalition is trying to compress that middle. If you are building AI native in europe and the bottleneck is compute plus access, this is one of the denser rooms on the continent right now.
Show more
Meet @nebius. They build cloud infrastructure for AI, including the GPU compute needed to train and run demanding AI workloads. We're putting that infrastructure in the hands of hackers for 48 hours. Barcelona brings the builders. Nebius brings the compute. Sept 19–20.
Show more
Two years ago Bernardo Kastrup was doing nights in an attic. Yesterday he announced a chip that consumes 100x energy maybe the next 🇪🇺 success story? dude is kind of a legend, ex-ASML, philosophy books, designed a chip aimed at neural nets that, on paper, burns a fraction of what a GPU burns because it stops dragging data back and forth across the package. This week his company raised €200m+, with Samsung as a co lead Peter Wennink (ex ASML CEO) is chair. they ship the physical system in 2028 (and license the silicon if someone else wants to build on it) Everyone agrees inference is becoming the meter. The technical problem is not mysterious, a lot of the watts never do the actual work. They pay for the trip between memory and compute. If you collapse that trip, you need fewer racks for the same token work, and basically that's the whole company. the Samsung link is quite interesting. Earlier on they were already in the manufacturing conversation, and now they are in the round. They already sell the memory, if compute and memory get designed together, that is their fight. Whether any of this works is still to see rn: no public rack, no measured watt/dollar, and until those exist, the performance claims are just claims but if that works, two things change. Cost per token stops being a GPU list price problem, and memory vendors stop being a side quest
Show more
if you could allocate all the compute in the world to one thing, what would it be?