Web agent economics are set by inference volume and each step is one inference call. A simple task (extract a field, fill a form, navigate a site) is 10 to 20 calls.
A multi-site workflow runs into the hundreds.
From
@RaiseSummit last month:
@abhshkdz,
@yutori_ai co-founder/CEO, on how they ship Navigator at 2-5x cheaper and faster than closed models on web tasks:
-> ~80% prefix cache hit rate
-> Speculative decoding on smaller footprint models
-> Post-training Qwen models with SFT and RL
They run on Together, where customizable endpoints and auto-scaling let them A/B test new model variants in minutes.