Register and share your invite link to earn from video plays and referrals.

DeepInfra
@DeepInfra
Fast ML inference. Run top AI models using a simple API.
76 Following    6K Followers
We just dropped support for newest models from @XiaomiMiMo The cheapest in the market, as always!
I used my entire usage limit for BOTH Codex and Claude Code (2 accounts) to tackle a local build, and the problem was never solved even though I used Fable 5.1 for intensive planning, parallel-agent supervision, and evals. I then went to @DeepInfra and got an API key for DeepSeek V4.1-Flash (US-hosted), and it is now CRUSHING this massive task Both Fable 5.1 & Astra could not. 50% of the way through & I only spent $1.15 so far. This is why all of a sudden the AI Cartel is sounding the alarm about AI safety. It's not about safety, it's about caping the competition. The numbers speak louder than perceived good will of the Effective Altruism mafia.
Show more
Encoder-decoder is back 😈!!! In DeepSeek V4.1 Flash the first 20 layers build the global KV that the next 20 read from, so prefill costs about half. Rolled out over the last 24h: throughput doubled, and we are approaching 1T tokens/day on OpenRouter. 30% off to celebrate. Enjoy. Cheapest on the market, as always.
Show more
Persimmon is a genuinely different idea: a model of how people actually talk, not another assistant. Proud to support @humansand on this launch with DeepCluster, a dedicated NVIDIA Blackwell cluster we deploy and operate. Excited to see where it goes.
Show more
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
Show more
0
87
1.4K
142
Forward to community
congrats to @humansand team on Persimmon release 💫
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
Show more
Ling-3.0-flash-VL is live on DeepInfra — day 0 with @AntLingAGI.
Today, we're open-sourcing Ling-3.0-flash-VL in BF16 and FP8. FP4 and INT4 are coming soon. Beyond visual recognition, it follows visual cues to: - Understand images, video, docs & UIs - Reason, search & verify - Use tools, check results & deliver
Show more
46% fewer errors 👀🚀
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: Academic Paper:
Show more
Congrats to @AntLingAGI on the launch of their new model!
We’re open-sourcing Ling-3.0-flash-Fin, a finance-enhanced model for real-world workflows, and FinFIRST, an expert-built benchmark for financial search agents. Two open releases, one goal: making financial AI more accessible and verifiable.
Show more
The @Zai_org team is on a roll. Another open-weight drop, and a seriously good one — congrats to everyone who shipped it. Live on DeepInfra now ->
GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: Tech blog:
Show more
Video and audio, from one model. Wan3.0-Video from @Alibaba_Wan is now on DeepInfra: 30-second clips at 1080P, omni-modal reference (images, files, even web pages), and characters that stay the same character. $0.20/sec →
Show more
congrats to @Zai_org team, big milestone!
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: Available now across all official platforms: Weights: API: Coding Plan: ZCode: Chat: AutoClaw:
Show more
THE RISE OF THE DATA CENTER Here's how much revenue Nvidia $NVDA has brought in each quarter over the last couple of years from its Data Center business Q4 2019: $1B Q1 2020: $1.1B Q2 2020: $1.8B Q3 2020: $1.9B Q4 2020: $1.9B Q1 2021: $2B Q2 2021: $2.4B Q3 2021: $2.9B Q4 2021: $3.3B Q1 2022: $3.8B Q2 2022: $3.8B Q3 2022: $3.8B Q4 2022: $3.6B Q1 2023: $4.3B Q2 2023: $10.3B Q3 2023: $14.5B Q4 2023: $18.4B Q1 2024: $22.6B Q2 2024: $26.3B Q3 2024: $30.8B Q4 2024: $35.6B Q1 2025: $39.1B Q2 2025: $41.1B Q3 2025: $51.2B Q4 2025: $62.3B Q1 2026: $75.2B
Show more
Our summer fastest-growing vendors list just dropped, and the story is: infrastructure won. Turns out the real AI hype isn't the app, it's what's under it.
New on DeepInfra: Qwen3.8-2.4T-A95B 🚀 @Alibaba_Qwen's latest sparse MoE — 2.4T total params, 95B active, 512 experts. Built for coding, agentic workflows, and complex reasoning, with native 262K context. Live now at $2.00/M in · $6.00/M out · $0.20/M cached @Alibaba_Qwen @alibaba_cloud
Show more
Ling-3.0-flash is live on DeepInfra A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN). Built for agents. Live now 👇
Show more
Step 3.7 Flash is Live on DeepInfra: An Agentic, Multimodal Model Built for Production
DeepInfra × Hugging Face DeepInfra is live on @HuggingFace Inference Providers. Run DeepSeek V4, Kimi-K2.6, GLM-5.1 and 100+ more open models straight from the Hub — same OpenAI-compatible API, same low per-token pricing, no markup. Just add :deepinfra to the model name.
Show more