SITUATION DETECTED: Nvidia is using its $6 billion deal with Poolside to develop one of the world’s most powerful open-weight models, aiming to challenge both Chinese labs and the US frontier. More than 100 Poolside engineers will be joining the Nemotron project, per WSJ.
we're seeing more traffic than ever and although most of our infra can scale flexibly it's not perfect
we're seeing some issues with large request bodies hitting up against some memory limits
continuing to work through them and improve things
Okay let me tell you about what's happening with DeepSeek v4 Flash.
First some background, it launched on Aug 1st and within 2 weeks it went from doing 3T tokens a day to 18T tokens a day on OpenCode; an absurd 6x increase. To put it in context, that's close to doubling up all of OpenRouter's daily volume. It's also likely 30-50% of DeepSeek's total volume.
Jumps like these over a 2-week span are not normal. This happened because the model is absurdly cheap; 350x cheaper than Fable, 175x cheaper than 5.6 Sol, 70x cheaper than Sonnet, and 7x cheaper than Luna. And secondly, it was a marked improvement over the previous Flash model.
For the first time our users got a feel for AI that's "too cheap to meter".
Then on August 16th DeepSeek raised prices by 5x (for peak hours, 2.5x off-peak hours). And it completely killed the growth of the model. It's doing less than half the tokens per day from its peak.
Obviously people were unhappy with the sudden change. Our guess is that DeepSeek genuinely could not handle the absurd 6x jump. Also, it's likely that the increase in GPU prices meant that even if they acquired new capacity, they wouldn't be able to serve it at the original prices.
A quick aside on why DeepSeek Flash is so cheap. It looks like they are running some custom infrastructure to cache way more tokens, for far longer. This matters because we've been scrambling trying to find providers that can fill this near 10T token per day void left by DeepSeek Flash. Unfortunately there are just a couple of people who are able to match DeepSeek's original pricing and that's likely only the case because they are using newer hardware.
That brings us to the current state of things. Over the last week we've talked to as many people as possible to get DeepSeek hosted at the original price. The issue is that even if somebody is able to, it's very hard for them to have enough capacity to handle our volume. It'll take roughly 1000 B300s to handle our throughput.
This is why if you've been using DeepSeek Flash on Go over the past few days, you might not have had the best experience. We've unfortunately cycled through a few different providers.
This 10T token per day gap, though, is an opportunity for every other model lab. It's very clear there's an appetite for a model that's at least as competent and cheap as DeepSeek Flash.
And somehow that still feels like the floor.
Ox Alpha (stealth model) is free for the next week
- 1M Context
- Multi-modal
- Zero Data Retention
Generous rate limits, near unlimited usage
We have capacity for 100T tokens per day, lets see what you can do
Ox Alpha (stealth model) is free for the next week
- 1M Context
- Multi-modal
- Zero Data Retention
Generous rate limits, near unlimited usage
We have capacity for 100T tokens per day, lets see what you can do
Muse Spark 1.2 Contributor is now available on OpenCode Go
this model will train on your data so it requires explicit opt-in and is restricted in some regions
in exchange you get incredible usage limits