Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
Joined July 2023
549 Following    11.2K Followers
๐Ÿคฏ Remember the weird U.S.-built diffusion LLM I posted about that broke 1,000 tok/s on Nvidia GPUs? It got a BIG upgrade. @_inception_ai released Mercury 2.5 today. ๐ŸŽฏ Instead of generating one token after another like a normal LLM, Mercury uses diffusion to generate/refine multiple tokens in parallel. And Inception says the new model delivers the following ๐Ÿ‘‡ ๐Ÿš€ 1,107 tok/s ๐Ÿ‘€ ๐Ÿง  +40% intelligence vs Mercury 2 ๐Ÿ“š 260K context ๐Ÿค” Tunable reasoning ๐Ÿ”ง Parallel tool calling ๐Ÿ“ Structured JSON ๐Ÿ’ต $0.20/M input / $0.75/M output ๐Ÿ‘€ ๐ŸŽฏInception says this is the largest diffusion language model ever trained. That ~1,100 tok/s isn't coming from Groq or Cerebras-style custom inference hardware - this is an important point! It's running on widely available NVIDIA GPU infrastructure. That's the big architectural bet, traditional LLM ๐Ÿ‘‡ token โ†’ token โ†’ token โ†’ token Diffusion LLM ๐Ÿ‘‡ many tokens โ†’ refine them together โ†’ answer โš ๏ธ Caveat, of course! ๐Ÿ”’ Mercury 2.5 is closed-weights.
Show more
Today, weโ€™re introducing Mercury 2.5, the most capable diffusion LLM on the market. It offers a 40% jump in intelligence over Mercury 2, and runs over 1,100 tokens/sec on widely-available @NVIDIAAI GPUs. Itโ€™s available today on our API, @OpenRouter, and @Baseten. Contact us to evaluate Mercury 2.5 for production:
Show more