Releasing our runtime dynamic compression framework that achieves 1.5-2.0 bit compression at high quality. This is integrated into the bitsandbytes2 library, which starts as a private beta today.
Paper:
Private beta signup:
Show more
Starting tomorrow, our lab will hold an open-source week: 2 software frameworks, 4 papers, all building a coherent ecosystem.
The theme: Frontier AI on Hardware You Own
Blog post:
SoTA results in:
Autocompaction
Autonomous Research
Model compression
Test-time scaling
Deep Research
And a healthcare RL environment giving you a new level of complexity to train healthcare agents.
The ecosystem that we will release is built to be as usable as possible. Agent sessions that run overnight and go on for tens of millions of tokens is made easy. Model compression of a model is automatic: you just get a good model that runs fast locally -- no expertise required. An autonomous research that works out of the box.
Our ecosystem enables a new level of work that can be done locally.
Show more
Have been using this model all day yesterday. With its new architecture, this release is just as big as R1 and will redefine all future models. It is such a big difference when you have a model close to 5.6 Sol, but it is 5 times faster and super cheap.
Show more
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.
What's new: 🥳
- Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4.
- Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks.
- Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
- 262K native context, extensible to 1M with YaRN.
We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀
We can't wait to see what you build with Qwen3.8-Flash!👀👇
- Blog:
- Technical Report:
- Hugging Face:
- ModelScope:
Show more
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog:
Available now across all official platforms:
Weights:
API:
Coding Plan:
ZCode:
Chat:
AutoClaw:
Show more
Real-time translation that sounds like you 🚀
We've been building speech-to-speech simultaneous translation for a while.
The newest piece is on-the-fly voice cloning: the model picks up your voice as you speak and carries it into the other language.
Here's the alpha running live during our all-hands: you speak in Japanese, the room hears you in English, and it still sounds like you.
It's early and the edges are rough, but this is the experience we've been building toward. Coming very soon.
Demo video below 👇
#
VoiceAI# #
SpeechAI# #
RealTimeTranslation# #
VoiceCloning# #
AI#
Show more
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.
Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.
Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Show more
One dead giveaway is also that Zhipu is one of the only labs that serves models with pretty poor partial prefill tok/s while output tok/s are fast. This model is ~30% faster than GLM 5.3. If the infras is the same, likely ~300-500B or fewer params vs activated params.
Show more
A new Kimi model, likely K3.1, is now being tested on the Code
@arena under the name "korrine"
K3 was tested on the Arena as "kivine" prior to its launch
If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu
Show more
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
Show more
🚀We’ve been pushing agentic inference toward the physical limits of the hardware.
Announcing LithosAI’s first pricing tiers, with early-access pricing ahead of the September 1 API launch.
Kimi K3 is live now at 800+ tokens/sec/user on standard GPUs, with full model quality.
Try the live demo at and sign up for early access.
Show more
I'm in Hyderabad next week, and I'd love to meet people outside my usual clique and bring them together.
Who are the best people in these areas that you know based out of Hyderabad: deep tech, AI product, AI research, GPU whisperers?
Show more
Compiler agent codesign is the next frontier in ML compilation and kernel agents
We’ve written an FAQ to answer some of the questions we've received about watermarking.
In summary:
• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
• Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs;
• The difference between watermarked and un-watermarked text will not be distinguishable to readers;
• Nothing is added to the text and there are no hidden characters;
• Watermarking doesn’t require extra tokens, and will not be more expensive;
• Watermarks can’t be traced to a specific person, organization, or chat.
Read more:
Show more
We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M tokens via YaRN.
- Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0.
🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently.
Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now!
Download, deploy, and build something we haven't imagined yet. 👀👇
- Hugging Face:
- ModelScope:
Show more
Trying GLM 5.3 right now. Avg: Prefill ~1 ktok/s, thinking/output ~60 tok/s.
In our harness, we already see with full thinking traces GLM 5.2 > Fable+Claude Code. But GLM 5.3 is just on another level. It is very precise and concise. Just testing long-task performance. Exciting!
Show more
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog:
Show more
With our new efficiency methods, you will be able to run this on a single DGX Spark or AMD Strix Halo at 7 token/s decode and >250 tok/s prefill. Stay tuned!
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog:
Show more
Fable quality. We will soon release a new quantized inference framework that will let you run this on a single GPU (B300) with ~50 tok/s in good quality. Local Fable LFG!
DeepSeek silently released V4-Pro 0813, up 15.8% on Terminal Bench from their April Preview model, with Fable 5 performance at ~57x cheaper cost.
1.6T param, 49B active, 1M context. This is the best price-to-perfomance model on the market right now.
Available in ClinePass now!
Show more
Businesses' spend on Fable 5 is almost the same spend as on Opus 4.6. Oof!
NEW from Ramp AI Index: disappointing adoption of Fable 5
We've heard several reasons from businesses...mainly Fable 5 is just too expensive.
A model so powerful it was briefly banned, and yet businesses don't think it's worth the price.
Show more
Our paper "XGBoost: A Scalable Tree Boosting System" received the Test of Time Award at #
KDD2026#! The honor belongs to the entire community that made
@XGBoostProject today. I will be giving a talk tomorrow at the award session on the paper, history, and lessons learned
Show more
anyone have recs for actually good articles about taste? not just developing it, but what it actually is