Register and share your invite link to earn from video plays and referrals.

Matei Zaharia
@matei_zaharia
1.5K Following    51.9K Followers
This is how Smart Routing working on our AI Gateway, super simple idea, lowers cost about 30% without giving up on quality!
Introducing Smart Routing in Unity AI Gateway. Stop overpaying for AI. Smart Routing matches each coding task to the right model and harness based on what the task needs, so higher-cost models can differentiate on intelligence while lower-cost models differentiate on cost and performance. Match frontier quality and cut task costs by 30%+.
Show more
I got this question so many times today. "How can you grow 80% at $7B?" The true answer is that we're finally seeing a breakthrough with AI agents starting to work in the enterprise. The AIs have been super smart for a while, but have lacked basic context that's in people's heads, or in some SaaS system-or-record. A lot of organizations are deploying FDEs to capture this context, or Ontology, and feed it to the AI. This is labor intensive and expensive. We just automated that with Genie Ontology. Once you have that enterprise context graph, an AI agent like Genie becomes magical. I find myself no longer waiting for answers from my CRO, CFO, CMO, CHRO etc, I just keep queuing up questions on the phone while sitting in meetings. It'd frankly addictive. Our customers are starting to do the same, over 70% of all queries on the platform are now generated by Genie agents. This fuels more questions to the platform, which drives consumption, which drives revenue. That's the simple answer.
Show more
0
80
1.2K
220
Forward to community
We are in a Cambrian explosion of frontier models. Even as a researcher who spends a lot of time evaluating models, I find it hard to keep up. New models are arriving constantly, and the mental overhead of choosing the right setup for each task keeps growing. That motivated us to build Smart Routing: an intelligent layer that automatically matches each task to the right model based on its complexity and the capabilities it needs. Our early results are promising: frontier level quality with 30%+ lower cost, and more than 50% savings on public benchmarks.
Show more
The fact that Databricks is growing nearly the same rate at $7b revenue as it was when the growth fund invested at ~$200m revenue 6 years ago never ceases to blow my mind. Today no one is better positioned to help the enterprise get true ROI out of AI. Congrats @alighodsi @rxin and the entire @databricks team! Honored to be your partners.
Show more
Smart Routing is now available in Unity AI Gateway on @databricks to improve coding agent quality and cost! Read how we built it to make it task-aware and preserve good cache hit rates. With so many frontier models coming out every week, this can really improve both cost and quality.
Show more
Introducing Smart Routing in Unity AI Gateway. Stop overpaying for AI. Smart Routing matches each coding task to the right model and harness based on what the task needs, so higher-cost models can differentiate on intelligence while lower-cost models differentiate on cost and performance. Match frontier quality and cut task costs by 30%+.
Show more
Introducing Smart Routing in Unity AI Gateway. Stop overpaying for AI. Smart Routing matches each coding task to the right model and harness based on what the task needs, so higher-cost models can differentiate on intelligence while lower-cost models differentiate on cost and performance. Match frontier quality and cut task costs by 30%+.
Show more
Today, we announced that we crossed $7B in revenue run-rate, growing over 80% year over year in Q2. We also shared: 🚀 $100M+ revenue run-rate for Lakebase 🚀 $1.5B+ revenue run-rate for Lakehouse, growing over 100% year over year 🚀 Continued positive adjusted free cash flow And we raised $5B in our latest fundraise. We’ll use this capital to invest in: 1️⃣ Lakebase, our serverless Postgres database built for AI agents 2️⃣ Genie, our AI coworkers that actually understand your business data 3️⃣ Unity AI Gateway, our multi-AI governance solution that helps control costs @iamVictorDey shares more in @Forbes:
Show more
We worked with @SpaceXAI to evaluate Grok 4.6 on the latest OfficeQA Pro V2 from @DbrxMosaicAI. It achieves the SOTA performance with @databricks's Genie harness! The model is strongest on our document understanding and data reasoning tasks, and it is a very efficient driver!
Show more
3 frontier models in one day! - Grok 4.6: Fable 5-level, but 85% cheaper. - Qwen3.8-Max: 2.4T params, 95B active. Weights are out. - DeepSeek-V4-Pro-0813: weights could drop any time. Heard it’s good, not just benchmaxxing. Competition is great for consumers and businesses.
Show more
0
67
2.3K
144
Forward to community
Omnigent v0.9.0 is now available! 🎉 🧭 Smart routing (@databricks AI Gateway + OSS) 🎨 Refreshed web UI 🧰 More sandbox and deployment control 🤖 Grok Build harness + example agents 🌐 nimble_extract and nimble_research builtins 🔗 Full details, bug fixes, and upgrade notes: #Omnigent# #AIAgents#
Show more
Did you know you can run Omnigent sessions straight from Slack? 👇 Mention the bot to start work, then approve tool calls and answer multiple-choice prompts inline in the thread. Learn more: #agents# #omnigent#
Show more
PGlite + real-time sync have emerged as key primitives in an era where millions of apps are deployed by agents. We’re excited to announce @ElectricSQL is joining team Neon at Databricks to build the world's most advanced Postgres backend platform.
Show more
Welcome to Databricks and @neondatabase, @ElectricSQL!
Excited to share that we've acquired ElectricSQL, the team behind PGlite. Agents need super fast Postgres and this team built an amazing WASM (WebAssembly) implementation of postgres that runs in your browser, but can sync back with Postgres instances asynchronously. Exactly what blazing fast AI agents today need. Excited to supercharge our 𝐋𝐚𝐤𝐞𝐛𝐚𝐬𝐞 𝐏𝐨𝐬𝐭𝐠𝐫𝐞𝐬 offering with these capabilities.
Show more
Excited to share that we've acquired ElectricSQL, the team behind PGlite. Agents need super fast Postgres and this team built an amazing WASM (WebAssembly) implementation of postgres that runs in your browser, but can sync back with Postgres instances asynchronously. Exactly what blazing fast AI agents today need. Excited to supercharge our 𝐋𝐚𝐤𝐞𝐛𝐚𝐬𝐞 𝐏𝐨𝐬𝐭𝐠𝐫𝐞𝐬 offering with these capabilities.
Show more
AI tokens are another resource to optimize in software engineering now. My cofounder @pwendell wrote about how we and other tech companies are starting to manage this resource now, by routing everything through an AI Gateway. This enables centralized analysis (e.g. we found settings we can change on Claude Code and Codex to lower cost a lot), smart routing, and “pushing down” control to our engineers so they can set budgets on individual tasks and prevent surprises.
Show more
Today @databricks we're publishing a detailed analysis of techniques we used to drastically reduce our internal AI spend while aggressively growing adoption. Savings come from layering in several techniques, which combine to drive unit costs down as much as 90% in some scenarios. Tl;dr, the wins come from: 1. Shifting defaults to more efficient models, including OSS models such as GLM. Maximum intelligence models simply aren't needed for many coding tasks, and "good enough" models are quickly becoming very cheap. We shift traffic between models using Unity AI Gateway. Approximate savings: 50% or more. 2. Using smart routing to automate model selection. Routing can further squeeze efficiency by dynamically selecting the model or harness that can most efficiently execute a particular task. Our task-level routing leverages @omnigent_ai. Approximate savings: 30%. 3. Providing user visibility and adaptive budgeting. Every user can see how much they spend, and users receive hints on how to contain spend. Heavy spenders encounter progressive friction as they ratchet spend above certain levels. Approximate savings: 10%. 4. Managing context bloat by pruning tool call results and tuning harness settings. Extraneous context costs $$ and delivers no value. Tuning cache settings also help lower average token costs. Approximate savings: 10%.
Show more
Today @databricks we're publishing a detailed analysis of techniques we used to drastically reduce our internal AI spend while aggressively growing adoption. Savings come from layering in several techniques, which combine to drive unit costs down as much as 90% in some scenarios. Tl;dr, the wins come from: 1. Shifting defaults to more efficient models, including OSS models such as GLM. Maximum intelligence models simply aren't needed for many coding tasks, and "good enough" models are quickly becoming very cheap. We shift traffic between models using Unity AI Gateway. Approximate savings: 50% or more. 2. Using smart routing to automate model selection. Routing can further squeeze efficiency by dynamically selecting the model or harness that can most efficiently execute a particular task. Our task-level routing leverages @omnigent_ai. Approximate savings: 30%. 3. Providing user visibility and adaptive budgeting. Every user can see how much they spend, and users receive hints on how to contain spend. Heavy spenders encounter progressive friction as they ratchet spend above certain levels. Approximate savings: 10%. 4. Managing context bloat by pruning tool call results and tuning harness settings. Extraneous context costs $$ and delivers no value. Tuning cache settings also help lower average token costs. Approximate savings: 10%.
Show more
0
71
1.8K
262
Forward to community
You can now branch your database AND your files without duplicating storage Neon Object Storage gives you S3-compatible buckets that branch just like your Neon Postgres does
I would like to not-at-all-humbly state that we were right all along (or at least since early 2024), not about anything neurosymbolic, but rather about how important it is to think in terms of AI systems, not just models:
Show more
📣 Excited to introduce OfficeQA Pro V2, the next generation of OfficeQA Pro! It's built on a new corpus of 120,000 PDFs provided by the U.S. Treasury. Frontier AI agents average 26% accuracy. More info below 👇
Show more
1/ We just made GEPA’s optimization loop much faster, along with improved generalization! Instead of proposing and evaluating one candidate at a time, each GEPA step now proposes a batch of candidates and evaluates them all concurrently. Parallel proposals significantly cut wall-clock time (3–4× faster) while, surprisingly, also generalizing better (up to +11 pts). 🧵
Show more