註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Praveen Neppalli
@praveenTweets
Chief Technology Officer @Uber
加入 May 2009
631 正在關注    12.6K 粉絲
As our CFO @_balaji_km mentioned at earnings today, we’re seeing some very interesting trends on AI costs. I think it’s another signal that we’re coming to the end of the so-called ‘tokenmaxxing’ era. Here’s what’s been happening behind the scenes. Since the beginning of the year we’ve more than quadrupled the number of people using frontier AI tools. That’s thousands of engineers using them every single day. During that same period, our cost per token has declined. You might expect costs to rise as adoption accelerates. We've seen the opposite. Not because we've restricted access, but because we've treated efficiency as an engineering problem rather than a budget problem. A few examples: • Caching and reuse: We use optimizations to improve our prompt cache hit rate that reduce our input token spend. • Better defaults and tooling: We tuned default model settings, context sizes and developer workflows so teams get the same results with fewer tokens and lower-cost inference. • Visibility drives efficiency: We gave engineers real-time visibility into their AI usage and costs per hour. • Experimenting with open-weight models: we continuously evaluate new models and deploy the best option for each use case. This is the future of applied AI at enterprise scale. The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible. Credit to all the engineers at @Uber who are helping to build this future. 🚀
顯示更多
0
76
1.6K
148
轉發到社區