가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Alvin Foo
@alvinfoo
Venture Partner | ex-Google | Decoding AI • Motivational stories on innovation & resilience • Tesla/SpaceX insights 🚀
가입 February 2007
12.2K 팔로잉 중    164.6K 팬
The economics are hard to argue with. When a model delivers strong and in many agentic/coding workloads competitive performance at a fraction of the price with dramatically better cache behavior and speed, high-volume usage shifts. We are already seeing this in the data: Chinese open-weight models (DeepSeek, Qwen, Xiaomi MiMo and others) now account for the majority of token volume on major neutral routers, around 60%+ versus US models in recent measurements, driven exactly by the cost and accessibility advantages highlighted here. This looks a lot like the Android story. Android never needed to win every premium benchmark or every enterprise RFP the way iOS did; it won by being open, cheap enough for the mass market, and good enough for the vast majority of real workloads. The result has been a durable ~70% global share. Chinese open-source models are following a similar path in the LLM space: flooding the open/inference layer where volume and price sensitivity dominate. 70-80% of the overall market is an aggressive but plausible ceiling for the open/commodity tier over the next couple of years if the current trajectory continues. Frontier closed models (Claude, GPT, Gemini) will keep the high-value, high-trust, multimodal, and heavily regulated slices, much like iOS keeps the high-margin end. But for the bulk of tokens processed in the world, the Android parallel is increasingly the right mental model. Cost and openness are winning the volume war.
더 보기