We post-trained a model for Harvey's Review Tables with
@appliedcompute, reducing costs by 50% while improving answer and citation quality.
Review Tables allow lawyers to upload up to 10,000 documents and ask up to 500 questions over each document - yielding up to 5 million model queries.
As a result they comprise a large fraction of our inference workloads. In July, the largest review table query on Harvey cost $26,000 in inference costs over a single run.
We partnered with
@appliedcompute to post-train GLM-5.2 to optimize both quality and cost for review table queries.
For training data, we started with a corpus of open-source legal data like filings, contracts and emails.
We then used frontier models to simulate different user personas and generate realistic review table queries over this corpus, alongside generating ground-truth answers to these queries sourced from the documents. Our applied legal research team did several rounds of review to ensure this data matched our quality bar.
With GLM-5.2 as a base, we post-trained a custom review table model in Harvey's review table harness on Applied Compute's AC2 platform.
This custom model beat all frontier models by 4-17% on answer quality and 11-19% on citation quality for review table queries.
It's also much cheaper: less than half the cost of Sonnet 5, and 1/10th the price of leading frontier models.
We also experimented with harness engineering alongside model training. We trained two Qwen 3.6-35B-A3B models in two separate harnesses: one using single-turn RAG and the other using agentic search. The agentic harness matched answer quality but drove down input tokens by 50% and output tokens down by 29%.
These results are promising early steps in improving quality and cost across Harvey using custom models.
Deep dive from
@vtrengarajan,
@nikogrupen,
@itsjuliopereyra,
@srice120, and Karl de la Roche at Harvey, and
@rhythmrg @caenopy @jacob_dphillips at Applied Compute: