instinct and muse are long-running agents. they are core examples of why demand for compute will just keep rising
we're in the era of long-running agents for regular people and not just for developers and people in the ai world.
we've been addicted to workflows that are super long-running and token-heavy, mostly in coding. regular people are still using chatbots, but that's starting to change (instinct is completely free right now and must be going through A LOT of compute)
others have tried to make "personal assistants" before but it seems like instinct and muse are the ones that will stick
a typical chatbot reply is 500 tokens whereas a long coding task can be millions of tokens because it plans things, it checks work, and iterates based on previous steps. when regular people actually use agents, there are going to be so many more companies that appear. selling compute would be a good place to be in the next few years.
at this point, even consumer companies might start owning their own gpus
Show more
instinct and muse are long-running agents. they are core examples of why demand for compute will just keep rising
we're in the era of long-running agents for regular people and not just for developers and people in the ai world.
we've been addicted to workflows that are super long-running and token-heavy, mostly in coding. regular people are still using chatbots, but that's starting to change (instinct is completely free right now and must be going through A LOT of compute)
others have tried to make "personal assistants" before but it seems like instinct and muse are the ones that will stick
a typical chatbot reply is 500 tokens whereas a long coding task can be millions of tokens because it plans things, it checks work, and iterates based on previous steps. when regular people actually use agents, there are going to be so many more companies that appear. selling compute would be a good place to be in the next few years. maybe we'll see even consumer companies starting owning their own gpus??
Show more
doesn’t it feel like every AI company is starting to do everything??
neoclouds are adding inference. inference companies are reserving or buying gpus. even routers (like openrouter and vercel ai gateway) need to guarantee gpu capacity to make sure requests get served.
an intuitive reason for this is everyone wants to make more money, but I think it's more nuanced. what everyone wants is to control where tokens go and under that is needing to control the gpu supply
- for neoclouds, they need to maximize the usage of gpus (so they add inference to do this)
- inference companies would never want someone else (like AWS or a neocloud) to be the reason why they can't serve all requests (so they reserve or own gpus so they're more independent)
- routers need compute providers with enough available gpus (or there’s nowhere to send requests)
when it gets harder to get compute (like it is right now) it makes sense for even routers to eventually reserve gpus themselves.
in the next few years, these labels are going to get outdated. I predict that many companies will be doing many of the same things
Show more
inference companies are taking on a lot of risk right now
they're signing longer contracts for compute because they never want to have customers and run out of compute at the same time. it'd be a big problem if users are growing, but they have to pause the product
it's hard to get compute fairly right now, so this is all they can do. inference companies are in a particularly tough spot because they work on a pay-as-you-go model, where customers pay per token and can leave whenever. clearly, this is a problem when you buy gpus on fixed contracts, because you'll pay for them even if customers don't use them
inference companies are simply stuck with no other choice. changing the way compute is bought will also make inference better!
Show more
all neoclouds rent out gpus, and gpus are the most expensive part. it feels like the other parts like the land, power, and building itself would cost more as they feel so large but for a 1 GW data center, it’s said that the gpu servers cost almost 2x as much as everything else combined
neoclouds have different ways to go about this. a pretty stark difference is the nebius model and the coreweave model
- coreweave doesn't own many data centers. they rent space. their customers have paid ~15–25% of what they owe for the contract upfront
- nebius owns some data centers and rents space too. ~70% of the deals it signed recently included upfront payments. for large deals, these payments covered a much larger amount (50–60%) of what it needed to build clusters
this comes down to needing to borrow $. nebius doesn't have to borrow much money because their customers can fund most of it. and, since they own land and power, customers are probably more comfortable paying (even if they know the cluster hasn't been built yet).
neoclouds have different financing models, yet clearly both nebius and coreweave are doing great!!
(though if demand stays this high, neoclouds may have to rethink how they do financing)
Show more
a TPU neocloud ecosystem is starting to happen. nvidia is the center of ai right now, and we see companies like etched try to compete on a chip level. google will compete to create an ecosystem of TPUs to compete with nvidia's neocloud ecosystem.
we don't give TPUs much attention but we already have two signs that they're definitely useful:
- last September,
@fluidstack became the first neocloud to deploy google TPUs. no longer in google's own cloud!
- a month later,
@AnthropicAI planned to use up to 1 million TPUs. and then this April, it agreed to use multiple gigawatts of next generation of TPUs
then, this May, blackstone agreed to create a new TPU cloud with google. a *5 billion* equity commitment!
Show more
I learned today that there are neoclouds with b300s that they (in a literal sense) can't plug in because they have no power
so are we compute-constrained, or are we really power-constrained?
I think this reminds me of when Satya from
@Microsoft said they had chips "sitting in inventory that I can't plug in"
right now I think we're a little bit of both, but power is so much slower to fix. there are grid issues for new grid connetions and equipment (like big transformers) have long waitlists.
there are ex-crypto energy people. it feels like they might be able to use the energy to their advantage. bitcoin miners have the power and grid connections, and many of course are already trying to pivot into AI 👀
Show more
follow the compute! for several of the biggest compute deals, the same group of gpus are behind them.
an ai lab needs compute. maybe it buys from a hyperscaler, the hyperscaler gets capacity from a neocloud, and the neocloud buys the gpus.
there are some different forms this can take like:
- labs buy directly from neoclouds
- hyperscalers buy from neoclouds and sell it to labs
- labs make agreements directly with chip companies
- chip companies can be backup customers to buy back capacity
I do think adding every deal that's announced can overestimate demand as sometimes there's overlap. if you don't think there's a bubble, these relationships between labs and providers will only get closer
Show more
clustermax is partly sensationalism, and it’s also a true effort that has provided transparency into the confusing market of compute providers. they do feel like the authority on compute without any “checks and balances” (for lack of a better phrase), because it's both the only thing out there, and it's everywhere
most people probably don’t find themselves in the weeds of the tiny details of everything (even like the reactions different groups have to clustermax), but this is the stuff I find fun so I’d like to share!
many things can be true at once.
1) clustermax has given customers more visibility (much needed). this one's obvious. now you can tell which providers are better than others by looking at a tier list. they run real workloads, and score them on things customers would care about (reliability, networking, etc)
2) the potentially biased dynamic of clustermax
- many have said that higher-ups at semianalysis hold equity in lots of companies they rank
- even if this didn’t factor into the testing, a few conversations with providers makes it clear that clustermax is on top of everyone’s minds.
- ask any compute provider and they might brush it off and say they don't care about clustermax, but you'll notice their mannerisms change as the conversation goes on. if they ranked highly, they might say something like "we love dylan, we're super close with him" if not, it's usually hesitation with something diplomatic (“we're working closely with the team")
3) incentive (mis)alignment to rank providers
- you can't rank every provider ever so it gives people insight to whoever's in the ranking
- you could bring in a third party to test providers but their incentive to get tested is to rise the clustermax rankings whereas your incentive might be to understand differences in a sea of compute (I'm unsure providers care about "standardizing" compute in the way everyone's talking about wanting to make compute fungible)
most would say clustermax a net positive (?). either way, clustermax is a cornerstone of compute and AI. if you believe that AI keeps expanding, companies on clustermax are only going to get bigger
Show more
Jon's giving me too much credit here. but yes if you work at a neocloud, compute provider, or anywhere adjacent in this newer landscape (market-making, commodities trading, financing), I could definitely use your mentorship 😊!
Show more
everyone's brokering compute right now. they're making a lot of $. we should talk about it
I always thought that tech hates middlemen. most people do (it's common sense. why pay someone for making an intro?). even freshman year me at Stanford knew this. for a pitch thing, my group’s entire idea was find a middleman and build something that removes them
brokering will stay for a while until the market becomes more efficient. thought the brokering model was only the referral fee model but there’s also marking up the provider’s price
- referral fee: introduce the buyer and provider, then get a % of deal
- markup: add to every gpu-hour and keep the difference throughout the contract
we live in really funny times bc you need to be careful if you want to do the markup model to not tell the buyer the provider's price or else you'll have to defer to the referral model and maybe a provider is mean and won't let you switch models HAHA
sometimes there’s a stigma around admitting your company brokers compute bc it feels like everything else in the company is fake. brokering def can't be the main thing but 1) compute/life runs on relationships 2) being in the middle gives you lots of valuable info
Show more
maybe VCs should lead the way in financing compute? they're experts in pooling investors through SPVs. financing compute for the mid-market could use something like this!
if you need a few gpus, you can just rent from a cloud or marketplace. if you need huge amount (10,000+), you'll work with providers and lenders
it's awkward when you're in between. people are calling this “mid-market compute”, when companies need enough gpus that public cloud gets expensive, but not enough to get the financing and custom terms of a huge buyer.
so much demand but neoclouds can't use all those contracts to borrow $ for gpus
why won’t investors and lenders just fund the neocloud directly? because they would have to trust that their customer's pay them, and the neocloud doesn't fail, and that they could pay them back
I feel as if the issue is investors don't want to finance an entire neoclouds (too much risk). in startup SPVs:
- investors pool their money
- one legal entity owns the startup shares
for compute, this SPV could own gpus and investors receive part of the revenue of the neocloud. this exists for huge deals (like
@IREN_Ltd's financing for their
@Microsoft contract). time to bring a smaller version to the mid-market??
Show more
inference is already making compute tradable. if you’re bearish on trading gpus, think about inference! my friend
@lihanc02 described inference providers as “partly compute traders” !!
they reserve compute, run models on it, then sell the output by the token. they make money by buying compute cheaply and getting more tokens from each gpu, but take the risk that some of the compute will not be busy
they take all different kinds of compute and turns them into something that's MUCH easier to compare (the same model’s tokens at a set price and speed)
you can already see this happening in the inference market:
- turn gpu capacity into tokens:
@togethercompute ,
@FireworksAI_HQ,
@DeepInfra,
@baseten
- turn their own chips into tokens:
@GroqLLC,
@cerebras,
@SambaNovaAI
- compare token price and speed:
@ArtificialAnlys
- compare and route requests between providers:
@OpenRouter,
@Vercel,
@RequestyAI
there’s no open market for inference yet, but lots of the pieces are already here
Show more
nvidia-smi is a good example. in any systems class, you learn pretty quickly that busy does not always mean useful. nvidia defines gpu-util as the % of time where at least one kernel was running. so 95% utilization does not mean the tensor cores were doing 95% of the math they could do.
highly recommend reading the DATE 2024 paper! they looked at TensorFlow workloads running on an A100. gpu-level utilization was high, but average instruction issue rate was below 50%. tensor-core instructions were below 5.2%.
all the newer gpu architecture changes make more sense - TMA handles more of the data movement instead of making threads do it, blackwell added TMEM so the accumulator doesn't take up so much of the register file, matrix multiply can also run asynchronously
tensor cores don't have to wait as much with those changes
Show more