Register and share your invite link to earn from video plays and referrals.

Gabe Pereyra
@gabepereyra
building @harvey with my bud @winstonweinberg
246 Following    12.9K Followers
When I shared @harvey’s model strategy a few months ago, there were two parts: 1. Build our own model 2. Use that to help customers do the same We’ve done the first. Now we’re hiring for the second: Harvey’s Private Model Program. We’re seeing huge demand from law firms to own their own intelligence leveraging private data. This will enable them to become frontier firms that get smarter with every client matter. You’ll:
 - Partner with law firms to build these systems. - Build the team that delivers this at scale.
 - Work with our technical org to define the platform that powers this team. We’re looking for a technical PM or founder type to own Harvey’s Private Model Program. DM me if this sounds interesting.
Show more
The hardest thing about building @harvey is doing what’s best for our customers despite immense pressure to do what’s easy. The easy thing would have been to force our customers onto consumption pricing before they were ready and serve them worse models to protect our margins. We chose to help our customers transition on a timeline that works for them and give them the best models in the meantime, even though it hurt our margins. This meant optimizing our product through routing, harness improvements, and post-training so we could serve frontier intelligence at an affordable price. It also meant building the infrastructure for customers to monitor and manage spend: usage dashboards, per-matter cost attribution, spend caps, and ROI reporting. As a result, we improved our gross margins from -50% to positive in a single quarter despite usage doubling month over month and continuing to serve the best models. Our philosophy is simple: do what’s best for our customers, even when it’s painful, hurts our margins, or draws criticism from competitors, X, and the press. We believe the most important part of building a company is earning and keeping your customers’ trust. You do that by doing the hard thing for them, even when it costs you.
Show more
Excited to try Jev in parts of our systems that need calibrated probabilities for categorical decisions. We currently use a hacky version of this idea: small LLM classifiers for routing, citations, parts of Vault, tool use, and user escalation. One challenge is that LLM softmax probabilities aren’t necessarily calibrated confidence estimates. It will be interesting to see how RLCD improves calibration over the naive approach. Jev doesn’t generate text, so its “hallucination-free” framing isn’t a full solution to hallucinations. But better routing, citation selection, and escalation could reduce hallucinations across the broader system. Longer term applications for law firms include matter selection, associate staffing, and predicting billing disputes. Also excited to see open-source implementation of RLCD so we can post-train these models ourselves.
Show more
The gold standard for evaluating complex legal work is partner review. Partners can cost over $3,000 an hour and associates $1,000, making expert review expensive at scale. Evaluating a model on LAB through expert review alone would cost millions of dollars. In practice, we combine rubric-based LLM scoring with sampled human preference. But rubrics miss errors beyond predefined criteria, and human reviewers get fatigued and make mistakes at scale. We built a generative reward model to bridge this gap by training agents to approximate partner review. We give these agents the original outputs, web search, and other tools to check our systems’ work. We find that their judgments correlate strongly with expert lawyer review. This approach will help us scale human review across model training and production products, and will be central to training Tenet 1.5.
Show more
Thrilled to be joining Harvey and working in an area that's so integral to society. AI is a transformative technology, but realizing its potential requires going deep in specific domains and working in close partnership with the professionals who live and breathe that work. Legal is an especially rich and consequential place to do that. Harvey's progress, combined with the rapidly maturing ecosystem of open-weight models and post-training infrastructure, creates an enormous opportunity to push the frontier and make AI transformative in practice. We're building out Harvey's founding research team and hiring across post-training skill sets for people who want to work at the intersection of research and product. If that's you, please reach out.
Show more
Excited to welcome @asadovsky as Harvey’s Chief Research Officer. Before Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As a CVP at Microsoft AI, he helped build MAI-Thinking-1, Microsoft’s reasoning model. As part of Gemini’s leadership team he helped train Gemini 1.0 through 2.5, including fine-tuning, RL, data, and evals. His prior work as a Distinguished Engineer at Google spanned Assistant, Search Quality, and Search Infrastructure. I met Adam three years ago when I sent him a cold LinkedIn DM and was surprised he responded. At a time when most dismissed the application layer and legal, Adam was curious and generous with his time. He quickly became someone I regularly turned to for advice on AI as we scaled Harvey over the past three years. When we first met, we were too early to hire someone of his caliber and scale, but I always hoped we’d eventually work together. As Winston and I got to know him better, what stood out even beyond his technical achievements was his character. Despite his incredible technical career, he remains curious, humble, practical, and cares deeply about the teams he builds. We couldn’t think of a better leader to help us build frontier intelligence for the professionals and institutions we serve.
Show more
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Show more
0
276
5.1K
525
Forward to community
We’ve raised $550M at a $15.5B valuation, led by @DiffusionVC and @LightSpeedVP. We’ve crossed $400M in ARR and serve 3,000 customers, including 80% of the top 100 law firms, 20% of the Fortune 500, and half of the Fortune 10. We are investing the capital from this round into our two most important resources: people and compute. Our product is expanding into new verticals and more specialized solutions for areas like contracting, litigation, deals, and compliance. As usage grows, we’re investing heavily in inference and model training to deliver frontier legal intelligence at the best possible price. Building on Tenet, our first model post-trained on open-weight models, this round will let us scale the compute and data behind future generations of our models. We believe the most successful application-layer companies will become full-stack AI companies, building across applications, agents, and models. Becoming a full-stack AI company will help our customers own more of their intelligence and build their vision of a frontier legal organization. We’re also grateful to @Sequoia, @KleinerPerkins , @A16Z , @CoatueMgmt, @Conviction, @EladGil, Evantic, GIC, @GoldmanSachs, Sapphire Ventures and Whale Rock for participating in the round and backing that vision.
Show more
We’ve raised $550M at a $15.5B valuation co-led by @lightspeedvp and @DiffusionVC. We're using this funding to help law firms, in-house legal teams, and professional services build and own their intelligence.
Show more
Awesome work by @nikogrupen @ItsJulioPereyra and team with @baseten. Together with @baselabs, we built a recursive language model (RLM) harness that lets a root agent delegate document review to sub-agents and combine their findings into a diligence memo. Post-training Qwen 3.5 in this RLM harness more than doubled rubric pass rate on LAB Diligence tasks, from 29.9% to 63.0%, and increased review coverage from 62% to 96% of all documents in a dataroom. We’re also doing an RL scale up run with GLM-5.3 and will share the results soon and believe model-harness co-optimization is a viable path to automating end-to-end legal tasks like M&A diligence.
Show more
Update on @harvey’s model training effort. We post-trained a model we are calling Tenet: - Achieves SOTA on LAB - Generalizes 3rd party legal benchmarks - Uses sub-agents for domain specific capabilities Tenet uses Kimi K3 as base and was post-trained in collaboration with @FireworksAI_HQ: - Rank-64 LoRA over the full network - GSPO with importance-ratio masking - 134 B300 GPUs for 2 months Despite not being trained on 3rd party legal datasets we found improvements on: - @mercor’s Apex Agents - Corporate Law - @crosbylegal’s Redline Bench - LegalBench Tenet also learned how to use domain-specific subagents (separate post-trained models) for complex tasks: - M&A Diligence: training in an RLM harness for long-horizon tasks (with @baseten) - Review Table: specialist models for high-volume structured data extraction (with @appliedcompute) - Firm Knowledge: parametric memory and structured notes for more efficient enterprise search (with @engram) These results suggest we can significantly scale training and we plan to: - Scale both human and synthetic data significantly and scale training to 1K and then 10K GPUs - This scale will let us move to full parameter fine-tuning and larger models - We are now starting to post train models in our production harnesses - Post-train other open-source base models to provide customers with model choice If these problems sound interesting we are hiring for our post-training team
Show more
Awesome research by @engramlab using the synthetic law firm we built. They trained a 27B Qwen model on synthetic client matters and used the learned knowledge to do online search more efficiently. The trained model outperforms frontier models at 10x lower cost per query. The results are promising for scaling up enterprise search and personalization.
Show more
Today we're publishing our first research blog, Understanding a Law Firm through Study. We're sharing a glimpse of a future where agents are trained with native memory:
In July, 20% of our inference spend was review tables. The most expensive review table queries cost $20k. Today we’re excited to share work we did with @appliedcompute that will reduce this cost by 50% by post training a review table model. Our review table product allows layers to upload thousands of contacts or emails and extract or analyze flexible fields. Historically we optimized this system via prompt caching, packing (multiple fields per model call), routing, batching, etc. Once we hit the limits of gradient free approaches we wanted to see if post training could improve on an already strong baseline. Review tables turned out to be a perfect candidate for post training because verification is easy and we’ve built large datasets to do so. Awesome work by @vtrengarajan, our vault and labs team as well as applied compute. Detailed write up below.
Show more
Had so much fun giving this talk at @sequoia about @harvey’s moneyball approach to building a research lab. The biggest mistake I made in the early days of Harvey was trying to play the Yankees baseball style of frontier intelligence. I found out the hard way that we were the Oakland As - we couldn’t raise the capital or attract the talent to build a frontier lab. However a lot has changed since then and it now feels possible to build frontier intelligence without a frontier budget. Winston and I’s most quoted line from Moneyball is “We can recreate him in the aggregate” when Billy Bean talks about his strategy for building the team The talk outlines our playbook to building frontier intelligence in the aggregate and how we leveraged the frontier ecosystem to do so. This is only now possible with inference providers like @FireworksAI_HQ and @baseten, neolabs like @trajectorylabs, @appliedcompute, and @EngramLab, data providers like @mercor, eval infra like @LangChain and many more. I talk about how we build training data and benchmarks, work with the neolabs and training infra providers to post-train, and give an overview of our serving and eval infra to ensure post trained models work in our product. At the end of Moneyball Billy says that if they don’t win everyone will dismiss this strategy but “if we win, with this budget, and this team, we will have changed the game”. Every application layer company, software company, frontier ecosystem company and startup now has a massive opportunity to play moneyball for frontier intelligence. Go change the game.
Show more
Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sovereign AI event, @gabepereyra shared @harvey ’s “moneyball” approach. Here’s the playbook: 00:00 Introduction 00:37 Building a research lab on a budget 02:28 Legal Agent Bench, contracting, and the diligence dataset 03:57 Domain experts guiding synthetic data generation 05:23 Why Harvey open sourced its datasets 06:55 Working with the neo labs – and why more than one 08:20 Post-training in-house: building "Associate 1" 09:44 The model serving matrix: 60 countries, fallbacks, SLAs 11:05 Deciding what stays in production 12:29 Simple open source switches and model routing 13:55 Moneyball: "If we win on this budget, we change the game" 14:53 Q&A: Training with sensitive data 17:16 Q&A: Competing for research talent 18:46 Q&A: Designing rubrics that actually challenge frontier models 20:19 Q&A: Where the pipeline breaks — data, research, or infra 22:59 Q&A: The tension in open sourcing a benchmark 25:02 Q&A: Biggest remaining open problems 27:10 Q&A: Competing with horizontal products
Show more
In collaboration with @harvey, we’re excited to build a new kind of agent environment to reflect realistic knowledge work: an entire synthetic law firm, Calderwood & Harkness, with over 100M (!) tokens of documents and 250 client cases. 💼 Today’s AI models know a lot about the law, but don’t understand how the law is practiced, because this information remains proprietary within firms. Unlike how agents are benchmarked today – starting each task from scratch with a new set of context — lawyers accumulate knowledge over time, building on years of experience. Calderwood & Harkness makes it possible for agents to do the same. Legal agents do many tasks in the same environment, making it possible to leverage memory and experience to do better work over time. We’ve had a great time co-developing this benchmark with Harvey’s research team, partnering @ItsJulioPereyra @nikogrupen @gabepereyra. We’ll share more results on this soon.
Show more
We are working with @EngramLab to train models that can understand a law firm’s knowledge. A huge source of differentiation for a law firm comes from all the past work the firm has done. When an associate is working on a new deal they have access to decades of similar deals the firm has done and a big part of the job is knowing how to effectively leverage that knowledge. However, a large amount of this data is either client data or derived from client data which means you can’t naively train all of it into a single model because this would mean breaking confidentially. In order to separate the research problem from the data privacy and deployment problem we built a synthetic law firm. @ItsJulioPereyra wrote an awesome article describing the dataset, tasks and approach to creating a synthetic version of the type of data you would expect to find at a large law firm. The firm has 46 clients, 266 client matters and each client matter contains work product, versions, emails, drafts, etc that you might find when searching the DMS of a firm. This dataset allows us to explore how well agents can reason across large corpuses of knowledge (100M+ tokens) and require complicated multi-hop reasoning. For example, answering a query like “in similar deals, how did we structure the reps and warranties” requires the model to understand what makes deals similar and then search over all past deals to find examples. We find that generic agent approaches for this type of reasoning are both expensive and not exhaustive and there is a lot of room for improvement. Excited to open source this dataset and also share more results soon
Show more
When I was an undergrad at USC, I emailed @JeffDean inviting him to come play soccer with the club team after his talk on campus. To my surprise he ended up moving a meeting with the dean to come and play with us. What is truly special about Jeff is that despite being responsible for some of the most impressive engineering and AI work in history, he still finds time for everyone - even if it means taking a break from inventing tensorflow to kick a ball around with an undergrad who looked up to him. I was fortunate to see some of the great work Jeff, @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix did at Brain and DeepMind. They are truly unique - so excited to see what this team can build.
Show more
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at:
Show more
Proud to have Goldman and JPM backing Harvey
Growth Equity at Goldman Sachs and J.P. Morgan's Growth Equity Partners are now investors in Harvey.
Was great chatting with @oneill_c and @mudithj about some of the research we've been doing together and broader AI topics
When @mudithj and I met @gabepereyra, we were expecting just another vanilla intro call and instead had the best yarn about research, the state of LLMs, and where intelligence is actually heading. It's rare to meet a founder this deep in the weeds who's also building for one of the most important verticals in this new age of intelligence So it was awesome to sit down with Gabe for an extended discussion on what it take to build agents that can reliably complete work over hours, days, or even longer? We talked about why agents today struggle with search and long context windows and how techniques like KV-cache compaction, synthetic data, and continual learning could help. 0:00 Introduction 0:36 Getting legal agents to review the whole data room 2:08 Data rooms larger than any context window 5:28 How far open-source models can go 7:58 Where specialist models fit in legal AI 10:59 Training legal models when client data is off-limits 13:06 Teaching a model how a law firm works 13:59 What belongs in context vs. model weights 15:36 From firm-wide AI to a model for every lawyer 18:37 What training adds beyond retrieving the right cases 20:26 Why context windows have plateaued 24:01 How models could learn continuously on the job 26:12 Can AI recursively improve AI research? 27:07 Research agents can run experiments but not choose them 30:00 Why open-ended research is hard to train 33:47 Why deployment, not intelligence, is the bottleneck 35:08 The cost of frontier intelligence 36:59 Different neolabs, different paths to intelligence 39:26 Using open datasets to compare research methods 41:13 Conclusion
Show more
We're open sourcing 10 diligence RL environments to expand @harvey's legal agent bench: - The largest one is 80M tokens - Human review of a dataroom costs ~$300K-$1M lawyer time - Each one has 100-1000 unit tests for automated evaluation This dataset will allow us to train diligence agents that can take an entire dataroom, client communication and deal context and product a diligence memo. Current frontier agents struggle due to dataroom size and domain knowledge around diligence but we are seeing improvements with post training. Awesome write up by @ItsJulioPereyra
Show more
Excited to welcome Alec, Connor and the Benchmark team to Harvey.