K3 already got in the top 5 most liked models of all time on Hugging Face, just 24 hours after being released! Ahead of Llama 3, Whisper and many other great models. Incredible!
Kimi K3 is now on Amazon Bedrock!
Run coding, document analysis, and extended agent workflows with Bedrock's access, encryption, and auditing controls. Explicit prompt caching supported.
Start building with K3 on AWS 👉
Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1.
We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results.
Thanks to the vLLM community for pushing Kimi K3 performance forward!
Read the deep dive:
While K3's usage figures have declined slightly in recent months, OpenRouter data currently shows as many as 300 billion tokens being generated each day by K3 models on the system.
Kimi K3 was released ~2 months ago, causing an uproar by the community due to its cyber capabilities being available in an open model
There has been no documented case of K3 causing harm so far
Kimi-K3 pricing update
Demand has grown faster than supply, so Kimi-K3 rates go up 30% on 2026-09-02 08:00 UTC.
Per 1M tokens:
• Input: $1.50 → $1.95
• Output: $7.50 → $9.75
• Cache read: $0.15 → $0.195