People keep asking why DeepSeek’s API is so cheap.
Some even make absurd claims that they’re dumping prices to corner the market.
No, the answer is simple: their model size is ridiculously small compared to its performance.
It's 10x smaller than Opus, and 5x smaller than Sonnet.
That means what used to require an 8-chip node can now run on a single chip.
And because it's so small, it runs extremely fast. You can multi-serve multiple users from a single chip while maintaining decent speeds.
By my math, they can handle 40x~ more traffic than Opus using the exact same compute.
This is where the competition is heading, and it’s why they can stay profitable even at these crazy prices.
Infinite money glitch :
> Borrow $1M from Japanese bank for ~3% APR
> Deposit into X money and get 6% APY
> You get $30K every year without spending your money
Prediction for Anthropic's next move:
They will warn that Chinese open-source AI has CCP backdoors and will steal your credit card passwords.
In response, the US government will impose strict restrictions on using Chinese open-source AI.
And society will start shaming people for using Chinese AI, all while your Claude freely harvests your data.
Just did the math and realized the inference hardware at my house now costs more than my car‘s used value.
2 DGXs, a maxed-out MacBook Pro, a Mac Studio, etc.
Honestly, I think it's the right choice.
I almost sold my Hyundai Genesis last time but backed out.
Now I really need to sell it and buy more hardware.
Cars depreciate, but hardware only goes up.
Anthropic could literally get sued by every human who has ever lived.
They just took all the knowledge humanity built up for thousands of years without paying a single cent.
At least Alibaba actually paid for the data they scraped.
The ultimate sovereignty starter pack:
- Mac Studio M5 Ultra 750GB
- Solar panels
- Power generator
- GLM-5.2
Building sovereignty right in your own backyard.
Chinese AI models have one massive advantage:
Their training data is mostly in Chinese, and a single Chinese character packs significantly more meaning than an English letter.
This means they can compress a lot more data per token.
It gives them up to a 4x token compression efficiency—an inherent advantage that English-based US models simply cannot replicate.
This is exactly why GLM, a mere 750B model, can compete with 2T-level frontier models.