๐คฏ Remember the weird U.S.-built diffusion LLM I posted about that broke 1,000 tok/s on Nvidia GPUs?
It got a BIG upgrade.
@_inception_ai released Mercury 2.5 today.
๐ฏ Instead of generating one token after another like a normal LLM, Mercury uses diffusion to generate/refine multiple tokens in parallel.
And Inception says the new model delivers the following ๐
๐ 1,107 tok/s ๐
๐ง +40% intelligence vs Mercury 2
๐ 260K context
๐ค Tunable reasoning
๐ง Parallel tool calling
๐ Structured JSON
๐ต $0.20/M input / $0.75/M output ๐
๐ฏInception says this is the largest diffusion language model ever trained.
That ~1,100 tok/s isn't coming from Groq or Cerebras-style custom inference hardware - this is an important point!
It's running on widely available NVIDIA GPU infrastructure.
That's the big architectural bet, traditional LLM ๐
token โ token โ token โ token
Diffusion LLM ๐
many tokens โ refine them together โ answer
โ ๏ธ Caveat, of course! ๐ Mercury 2.5 is closed-weights.