We've been investing heavily in our harness and routing capabilities, and we put them to the test benchmarking Glean's token costs against Claude Cowork. The results were striking:
@Glean is 4x more cost-effective, averaging $0.45 per task versus $1.84 for Claude Cowork.
That 4x advantage comes from two things compounding: 2.9x lower token volume and a 1.4x cheaper blended rate per million tokens. Here's how:
- Model family routing: Glean made use of Luna which is 10x cheaper than Claude Sonnet and widely capable. We’re able to strike the balance by routing between open and closed models.
- Model tier routing: In Glean, Opus was used 10x more (29% vs 2.8%) but surgically for the right things and balanced by other models.
- Better context: Glean’s harness and indexing capabilities result in fewer tokens consumed; Claude Cowork used 3x the tokens per query on average using 88.8M versus 29.8M in Glean.
More results coming out at Glean:GO! Hit me up if you're still looking for an invite.