登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Matei Zaharia
@matei_zaharia
参加 October 2010
1.5K フォロー中    52.5K ファン
AI tokens are another resource to optimize in software engineering now. My cofounder @pwendell wrote about how we and other tech companies are starting to manage this resource now, by routing everything through an AI Gateway. This enables centralized analysis (e.g. we found settings we can change on Claude Code and Codex to lower cost a lot), smart routing, and “pushing down” control to our engineers so they can set budgets on individual tasks and prevent surprises.
もっと見る
Today @databricks we're publishing a detailed analysis of techniques we used to drastically reduce our internal AI spend while aggressively growing adoption. Savings come from layering in several techniques, which combine to drive unit costs down as much as 90% in some scenarios. Tl;dr, the wins come from: 1. Shifting defaults to more efficient models, including OSS models such as GLM. Maximum intelligence models simply aren't needed for many coding tasks, and "good enough" models are quickly becoming very cheap. We shift traffic between models using Unity AI Gateway. Approximate savings: 50% or more. 2. Using smart routing to automate model selection. Routing can further squeeze efficiency by dynamically selecting the model or harness that can most efficiently execute a particular task. Our task-level routing leverages @omnigent_ai. Approximate savings: 30%. 3. Providing user visibility and adaptive budgeting. Every user can see how much they spend, and users receive hints on how to contain spend. Heavy spenders encounter progressive friction as they ratchet spend above certain levels. Approximate savings: 10%. 4. Managing context bloat by pruning tool call results and tuning harness settings. Extraneous context costs $$ and delivers no value. Tuning cache settings also help lower average token costs. Approximate savings: 10%.
もっと見る