Imagine doing your job without ever looking anything up on the internet. That's an agent without web search.
Giving agents the web changed what they can do, especially on anything recent that isn't reflected in training data. But agents don't search like people, so optimizing their performance is a new challenge.
There's a lot of great research out there on search behavior, but we wanted to answer a more operational question: when should search be on, and how should you configure it?
We evaluated 1,329 current events questions across 4 models and 14 conditions, comparing
@youdotcom, provider built-in search, and no search. We found that:
- Search reduced the gap between models from 47.9 points to 5.6
- Retrieval gain declined with event age, from ~45 points for recent events to ~24 for the oldest
- Runs with 5+ searches scored 19–49%. A fifth query was associated with lower performance
Read the research →