登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Mars_DeFi
@Mars_DeFi
Researcher and Visual Educator : DeFi • AI • NeoFinance • Tokenization | Top Research in Highlights | Channel
参加 December 2014
3K フォロー中    27.5K ファン
The next AI bottleneck may not be better models. It may be the physical infrastructure needed to run them. AI is shifting from a world dominated by massive training runs toward one where inference becomes the larger, recurring source of compute demand. And AI agents accelerate that shift. A chatbot might answer a prompt once. An agent can reason, call tools, retry tasks, search, write code, and keep operating in the background of an application. That can mean 5–50x more token consumption, helping push total compute demand toward a potential 24x increase by 2030. So the AI boom is increasingly becoming a power, data-center, chip, and financing problem. — ● AI is becoming physical infrastructure AI feels like software because the output is digital. But producing that output depends on a very physical stack. Electricity -> grid -> data center -> GPUs + HBM -> AI cloud -> model -> application / agent Every layer depends on the one below it. • A new model is not very useful without enough compute to serve it. • More GPUs do not help if the data center cannot get enough power. • More power does not help if grid connections, transformers, cooling, or permits are delayed. That creates a simple relationship: More AI usage -> more compute -> more power -> more infrastructure Compute is increasingly becoming the raw material behind machine intelligence. — ● The AI economy can be thought of in three layers At the bottom is compute. Data centers, GPUs, networking, memory, and AI clouds convert physical infrastructure and energy into computation. Above that sit models. Proprietary and open-weight models turn computation into intelligence. Then come applications and agents. They turn that intelligence into actual economic activity by writing code, analyzing data, executing workflows, answering customers, and increasingly acting autonomously. So the stack looks like: Infrastructure -> compute -> intelligence -> action And demand flows backward through it. • More agents create more inference. • More inference requires more compute. • More compute requires more physical capacity. — ● Inference changes the economics Training gets most of the attention because individual training runs can consume enormous amounts of compute. But training is episodic and inference is continuous. Every time a user sends a prompt, an application calls a model, or an agent completes a task, compute is consumed again. Training still represents roughly half of current AI compute demand, but falling model costs are making inference much easier to scale. Pricing from models such as DeepSeek and OpenAI continues to push the cost per million tokens lower. That does not necessarily reduce total compute demand rather it can do the opposite. Cheaper tokens -> more applications -> more usage -> more agents -> more inference -> more compute This is the same dynamic seen across other technologies. When the unit cost falls enough, usage expands faster than the savings. — ● Agents amplify that effect This is where the infrastructure thesis gets more interesting. A traditional chatbot interaction is relatively simple. One prompt comes in and one response comes out. Agents can create much longer execution loops. They can: • Reason across multiple steps • Search external sources • Call APIs • Generate and execute code • Review their own output • Retry failed tasks • Coordinate with other agents Each additional step consumes tokens and compute. So if AI shifts from people occasionally chatting with models toward software continuously using agents, inference becomes a much more persistent infrastructure load. That is why future compute demand may be driven less by a handful of giant training runs and more by billions of small, continuous inference workloads. — ● The problem is that compute supply has bottlenecks everywhere AI demand can grow almost instantly but physical infrastructure cannot. Also new capacity has to pass through several constraints. • At the power and grid layer, utilities need enough generation and transmission capacity to support new loads. • At the data-center layer, projects need land, permits, transformers, cooling systems, and years of construction. • At the hardware layer, supply depends on GPUs, HBM, networking equipment, and advanced semiconductor packaging. • And at the political layer, projects increasingly face questions around power consumption, water, land use, and grid reliability. So a company may have the GPUs, customers, and capital and still be unable to deploy compute because one physical constraint has not been solved. That is why energized capacity matters more than announced capacity. — ● Private capital is moving directly toward those bottlenecks A growing group of infrastructure companies is attacking individual constraints across the stack. • @Etched is focused on specialized inference chips and has raised roughly $700M at a reported $21B valuation. • @_panthalassa is working on alternative energy infrastructure and has raised around $140M at a valuation approaching $1B. • VEIR is focused on increasing power-delivery capacity and has raised roughly $75M at a $170M valuation. • Madrone is developing advanced cooling systems aimed at reducing the power and water required by data centers. • @valaratomics is pursuing nuclear power infrastructure and has raised roughly $1B at a $6B valuation. These companies may look very different but they are all attacking the same underlying problem: AI demand is growing faster than physical compute capacity can be deployed. — ● Crypto is building another route into the same market Decentralized compute is approaching the bottleneck from a different direction. Instead of relying entirely on hyperscalers and vertically integrated neoclouds, crypto networks can aggregate fragmented capacity, create compute markets, or financialize the hardware itself. @bittensor coordinates decentralized AI services and compute. @akashnet operates a decentralized marketplace for GPU capacity. Venice and DIEM (@AskVenice) are experimenting with tokenized inference capacity. @USDai_Official is taking the financing angle, using GPUs as part of an onchain credit model. So the two infrastructure stacks begin to look like: • Centralized Hyperscalers -> data centers -> neoclouds -> AI workloads • Decentralized Crypto networks -> GPU markets -> tokenized compute -> onchain financing They are not necessarily replacing one another. They are different ways of allocating and financing scarce compute. — And that is the bigger shift happening underneath the AI boom. The market has spent years focusing on who can build the smartest model. But as models become cheaper and agents become more widely deployed, the strategic bottleneck moves downward. It becomes a question of who can supply enough Power, Grid access, Data-center capacity, GPUs, Memory, Cooling and Financing. Inference turns AI from a one-off training problem into a recurring infrastructure problem and agents make that demand even more persistent. So the biggest opportunity may not sit only in intelligence itself. It may sit in the physical and financial infrastructure required to deliver that intelligence at scale.
もっと見る