登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Gabe Pereyra
@gabepereyra
building @harvey with my bud @winstonweinberg
参加 February 2022
246 フォロー中    12.9K ファン
Had so much fun giving this talk at @sequoia about @harvey’s moneyball approach to building a research lab. The biggest mistake I made in the early days of Harvey was trying to play the Yankees baseball style of frontier intelligence. I found out the hard way that we were the Oakland As - we couldn’t raise the capital or attract the talent to build a frontier lab. However a lot has changed since then and it now feels possible to build frontier intelligence without a frontier budget. Winston and I’s most quoted line from Moneyball is “We can recreate him in the aggregate” when Billy Bean talks about his strategy for building the team The talk outlines our playbook to building frontier intelligence in the aggregate and how we leveraged the frontier ecosystem to do so. This is only now possible with inference providers like @FireworksAI_HQ and @baseten, neolabs like @trajectorylabs, @appliedcompute, and @EngramLab, data providers like @mercor, eval infra like @LangChain and many more. I talk about how we build training data and benchmarks, work with the neolabs and training infra providers to post-train, and give an overview of our serving and eval infra to ensure post trained models work in our product. At the end of Moneyball Billy says that if they don’t win everyone will dismiss this strategy but “if we win, with this budget, and this team, we will have changed the game”. Every application layer company, software company, frontier ecosystem company and startup now has a massive opportunity to play moneyball for frontier intelligence. Go change the game.
もっと見る
Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sovereign AI event, @gabepereyra shared @harvey ’s “moneyball” approach. Here’s the playbook: 00:00 Introduction 00:37 Building a research lab on a budget 02:28 Legal Agent Bench, contracting, and the diligence dataset 03:57 Domain experts guiding synthetic data generation 05:23 Why Harvey open sourced its datasets 06:55 Working with the neo labs – and why more than one 08:20 Post-training in-house: building "Associate 1" 09:44 The model serving matrix: 60 countries, fallbacks, SLAs 11:05 Deciding what stays in production 12:29 Simple open source switches and model routing 13:55 Moneyball: "If we win on this budget, we change the game" 14:53 Q&A: Training with sensitive data 17:16 Q&A: Competing for research talent 18:46 Q&A: Designing rubrics that actually challenge frontier models 20:19 Q&A: Where the pipeline breaks — data, research, or infra 22:59 Q&A: The tension in open sourcing a benchmark 25:02 Q&A: Biggest remaining open problems 27:10 Q&A: Competing with horizontal products
もっと見る