Register and share your invite link to earn from video plays and referrals.

Aaron Levie
@levie
ceo @box - your business lives in content. unleash it with AI
Joined March 2007
952 Following    3.5M Followers
Agents already make up the majority of inference. This will quickly trend toward nearly all inference over the next year or two. The vast majority of tokens used in the world will be agents that are executing unbelievable amounts of tasks for us in the background 24/7. Agents will be deployed to read all code changes to secure our software, process all of our data inside of workflows, handle a significant majority of the research that goes into recruiting and customer prospecting, review every event stream and log from every system in the world, go out and execute tasks for us in our personal lives, and hundreds of other use cases. The rate of new agents coming online that will be consuming insane amounts of tokens is not slowing down. Just in the past week I’ve introduced multiple completely new workflows that would not have been possible technically even a month ago. Incredible time to be doing anything in inference and of course building on top of all of this.
Show more
AGENTIC TRAFFIC NOW MAKES UP MORE THAN 70% OF ALL INFERENCE TRAFFIC 🚀 Agentic workloads are characterized by four elements: 🟠 Multi-turn: a session includes tens or hundreds of turns, leading to high potential KV-cache reuse. 🟠 Long context: system prompts, tool definitions, and the large number of turns make context accumulate quickly. 🟠 High prefix reuse: since the conversation progresses linearly, where output from turn n-1 is concatenated to turn n (typically), most context can be served from KV cache rather than recomputed (this depends on the amount of storage available to store KV tensors). As n grows, the ratio of cached input relative to uncached input typically tends towards 1. 🟠 Sub-agent bursts: a session launches multiple short-lived sub-agents with fresh context, which create bursty KV-cache patterns.
Show more