Register and share your invite link to earn from video plays and referrals.

vCluster
@vclusterlabs
Create & Manage Tenant Clusters Like a Hyperscaler
23 Following    3.7K Followers
@LukasGentele sat down with @dylan522p at AI Infra Summit. They got into why an H100 is worth more today than the day you bought it, even as token prices fall by an order of magnitude. Cost per token keeps collapsing, but the volume of tokens coming out of each GPU rises faster than the price falls, so total dollars per GPU keep going up. Compute gets more valuable, not less. Dylan then walked through what SemiAnalysis found testing AI cloud security. Many of the vulnerabilities could be surfaced just by pointing a model at the cluster and asking it to look. Other tenants' storage, other tenants' nodes, other tenants' Slurm. On one cluster they found a country's national security apparatus, multiple agencies, sitting in the same data center they had just gotten into. His summary: it's the Wild West. Which raises an uncomfortable question for every sovereign AI buildout currently being announced. They also covered why the plan almost every AI cloud starts with, hire a team and build the whole software stack ourselves, does not survive contact with reality. The hyperscalers absorbed the talent. The biggest AI clouds took the rest. As Dylan put it, the talent does not exist, and it takes far more people than anyone budgets for to build it, stress test it, and prove it has no security holes. The question they didn't resolve: if security is what pushes large customers into renting entire sites rather than sharing infrastructure, does the middle of the market ever get to exist? Or does everything consolidate into a handful of operators big enough to sell trust along with capacity? 0:00:00 - Why an H100 is worth more now than three years ago 0:01:49 - What an AI cloud can actually optimize in the stack 0:04:15 - ClusterMAX, and what separates gold tier from silver 0:06:19 - Point a model at a cluster and it finds the holes 0:09:48 - Bare metal off take deals and billion dollar rentals 0:11:23 - Where the bottleneck moves in 2027 0:13:05 - "The talent does not exist" 0:15:47 - What the AI clouds still standing will be selling 0:16:19 - Single tenant inference vs public endpoints
Show more
If you're at AI Infra Summit, this is the panel you shouldn't miss! The Speed Problem: What It Actually Takes to Stand Up Capacity ⏰ Wed, Sept 16, 3:00 PM PT 📍Compute Track at AI Infra Summit in Santa Clara vCluster CEO Lukas Gentele is moderating this one live with Jay Jubran (@ZyphraAI), Kasra Danesh (@sfcompute), and Yujing Qian (@gmi_cloud) Three companies. Three completely different bets on the same problem. Zyphra trained a frontier mixture-of-experts model end to end on AMD silicon, zero NVIDIA in the stack, then turned that build into an AI Cloud. SF Compute built a physically settled compute market. Long-term contracts get supercomputers financed. Resale lets customers recover an average of 25% of what they spend on capacity they don't use. They just moved $245M of Blackwell B300 across two deals. GMI Cloud stayed on the default stack and bet everything on speed: NVIDIA Reference Architecture, a 99.9% uptime SLA, $500M in new capex behind nine-figure enterprise contracts. Silicon. Capital. Software. What's actually standing between "we announced capacity" and "a workload is running on it,". Let's find out. #AIInfraSummit# #AIInfrastructure# #AICloud# #Compute# #Inference#
Show more