I heard about Not Diamond when learning OpenRouter since we pivoted towards inference as a service, excited to help them announce a new product optimizing long agentic workloads instead of per-request
Basically if you just route to the cheapest model, it can loop and produce bad work a strong model has to repair later, so you pay more. Optimize for merged code, not token consumption only
Also learned from docs that Not Diamond Code is cache-aware. Switching models mid-session wipes your warm cache and re-reading the entire context uncached can cost more than just staying on the expensive model. Especially important since github etc. are pivoting to multi agent across different parts of the workflow
Team said to expect 20%+ lower inference cost so we'll see! I'm continuing to explore demand side of AI token factory and plan to post more, mainly questions ๐
Today weโre announcing Not Diamond Code, the worldโs most powerful intelligent model router for long-horizon coding agents.
Not Diamond works with any gateway or harness, including Claude Code, to select the best model and reasoning effort for each step, reducing costs by 20-65% without impacting quality.