Register and share your invite link to earn from video plays and referrals.

Dan Biderman
@dan_biderman
1.6K Following    3.6K Followers
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here:
Show more
0
72
406
115
Forward to community
great companies here. there's a unique window to do stuff now
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Tacit Labs (@ninklefitz, @AmDroste) - American Terawatt (@atroyn, @rslparker, @aranibatta) - Conduit (@clemvonstengel, @riopopper) - Convergent (Omkar Savant, Vivek Katara, @debnilsur) - Core Automation (@MillionInt, @_arohan_) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Instinct (@noahrshinn) - Keenable (@styskin, Matthias Petri) - Lumaril (Mark Elliot, Ben Duffield) - Neion Bio (@Dimkell, Sam Levin) - Pangram Labs (@max_spero_, @bradley_emi) - Quadrillion (@echinaceous) - Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh) - Ricursive (@annadgoldie, @Azaliamirh) - Sail Research (@neilmovva, @blintzbase) - Trajectory (@rronak_, @michaelelabd, @QuantumArjun) - Watney Robotics (Sean Cheong, Ryan Gannon) Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg. Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote. The full list is on Breakout List.
Show more
Awesome news for everyone involved
Excited to welcome @asadovsky as Harvey’s Chief Research Officer. Before Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As a CVP at Microsoft AI, he helped build MAI-Thinking-1, Microsoft’s reasoning model. As part of Gemini’s leadership team he helped train Gemini 1.0 through 2.5, including fine-tuning, RL, data, and evals. His prior work as a Distinguished Engineer at Google spanned Assistant, Search Quality, and Search Infrastructure. I met Adam three years ago when I sent him a cold LinkedIn DM and was surprised he responded. At a time when most dismissed the application layer and legal, Adam was curious and generous with his time. He quickly became someone I regularly turned to for advice on AI as we scaled Harvey over the past three years. When we first met, we were too early to hire someone of his caliber and scale, but I always hoped we’d eventually work together. As Winston and I got to know him better, what stood out even beyond his technical achievements was his character. Despite his incredible technical career, he remains curious, humble, practical, and cares deeply about the teams he builds. We couldn’t think of a better leader to help us build frontier intelligence for the professionals and institutions we serve.
Show more
Underrated interpretability result from our work with @EngramLab on @harvey law firm knowledge that makes a difference for the UX of long-horizon agents. The reasoning traces you get from parametric memory + study notes are much more intuitive than the interleaved bash commands you get from coding agents. When you read them (esp as a non-technical user), you can actually understand what’s going on. For example: "I remember two antitrust matters from earlier work, one that was terminated (client matter 1001-00004) and one that cleared after an HSR Second Request (client matter 1003-00003)" Compare this to the following from a coding agent (1 of 19 such bash commands btw): { "command": "cd /tmp/cache/1003-00003 && echo \"=== engagement letter head ===\" && sed -n '1,25p' Engagement/engagement-letter-hpe-fund-iv.docx.txt | cut -c1-500 && echo && echo \"=== practice group / partner mentions ===\" && grep -rhoi \"[^.]*\\(practice group\\|M&A partner\\|antitrust partner\\|responsible partner\\|lead partner\\)[^.]*\\.\" . | sort -u | head -12 | cut -c1-300 && echo && echo \"=== HSR outcome ===\" && grep -rhoi \"[^.]*\\(waiting period expired\\|expiration of the \\(extended \\)\\?waiting period\\|early termination\\|closed the investigation\\|closing letter\\|consent decree\\|no further action\\|without taking\\)[^.]*\\.\" . | sort -u | head -15 | cut -c1-400" } Interpretability like this becoming increasingly important for enterprise agent deployments. s/o to @dan_biderman @realJessyLin & team for innovating on multiple dimensions here.
Show more
human-human communication is one of the last things to automate with AI declaude your emails
GLM 👏 Gemini-3.7 Flash also deserves 👏 for that price
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: Available now across all official platforms: Weights: API: Coding Plan: ZCode: Chat: AutoClaw:
Show more
awesome work on this new platform!
Today, we're introducing AC2, the Applied Compute Agent Cloud, to enable every team to train, serve, and improve their own frontier models.
Should be fun
Today, we're opening up our 4th annual Fellows Forum for registrations. We will have ~500 industry leaders gathered for two amazing days. Hosted alongside @nebiusai and @nvidia. Speakers include (but not limited to)... - @peteflorence (Generalist) - @LiamFedus (Periodic Labs) - @MillionInt (Core Automation) - @BorisMPower (OpenAI) - @edchi (DeepMind) - @liu_mingyu (NVIDIA) - @ShivdevRao (Abridge) - @pirroh (Replit) - @marcboroditsky (Nebius) - @JasonMa2020 (Dyna Robotics) - @dan_biderman (Engram) - @BanghuaZ (RadixArk) - @zacharylipton (CMU) - @jingli9111 (Stealth) Register here:
Show more
Very ambitious! Excited to bring some new capabilities and experiences to lawyers.
Rox has been powering Global 2000 enterprises with $5T+ in combined market cap. Today, we’re putting it in everyone’s hands. Over the last 2 years, companies like @MongoDB, @togethercompute, and @Xbow shifted investment from legacy CRM and SaaS to Revenue Agents. But most teams have been locked out. They lacked in-house technical talent or FDEs required to set up the revenue-specific context, harnesses, and agent systems needed to run Revenue Agents in production. Today, we launched Rox Teams to remove that barrier. With Rox Teams, businesses of every shape, size, and vertical can activate Revenue Agents on their own. Set up in minutes. No FDEs required. & Nolaned this hilarious explainer video 👇
Show more
0
189
1.4K
417
Forward to community
i'm quite excited about this! • it's cool in general that models can generate their own training data and learn from it. wasn't fully clear to me even a year ago that this would work reliably. • the thing that we're trying to build seems really important and no one has built it before • some traces from our model genuinely surprise me (e.g. the attached example, where it perfectly simulates the output of a complicated bash query, despite not being trained to do this) my main feeling overall is that there is so much to do. we see signs things are starting to work: models use their memories to produce better answers, knowing things saves lots of tokens, and all of this emerges with more compute but solving this memory calibration problem, on top of learning how to generate data in ways that scale nicely with compute, is going to take some time this is just a first step 🫡
Show more
Like real people, Engram's models use both neural memories and written notes to tackle hard tasks. Many legal tasks require reingesting the same persistent context: the firm filesystem. Scaling compute preemptively on that filesystem can obviate many challenges at inference time and unlock new types of behaviors.
Show more
Today we're publishing our first research blog, Understanding a Law Firm through Study. We're sharing a glimpse of a future where agents are trained with native memory:
working with @EngramLab to train models on law firm knowledge, we built a synthetic law firm (46 clients, 266 matters, 100M+ tokens) to study how well agents can search and reason across a firm's entire body of work.  more results to share soon!
Show more
Harvey Specter has always been a hero of mine and today I got one step closer to maybe asking @gabepereyra @winstonweinberg for an intro.
In collaboration with @harvey, we’re excited to build a new kind of agent environment to reflect realistic knowledge work: an entire synthetic law firm, Calderwood & Harkness, with over 100M (!) tokens of documents and 250 client cases. 💼 Today’s AI models know a lot about the law, but don’t understand how the law is practiced, because this information remains proprietary within firms. Unlike how agents are benchmarked today – starting each task from scratch with a new set of context — lawyers accumulate knowledge over time, building on years of experience. Calderwood & Harkness makes it possible for agents to do the same. Legal agents do many tasks in the same environment, making it possible to leverage memory and experience to do better work over time. We’ve had a great time co-developing this benchmark with Harvey’s research team, partnering @ItsJulioPereyra @nikogrupen @gabepereyra. We’ll share more results on this soon.
Show more
Frontier models know law. But there are many ways to practice law. Everything we know about them can only be found in the huge filesystems of law firms. To me, how models understand and navigate those mega filesystems is the big open question in AI.
Show more
Excited for them and more soon
Introducing Harvey Research: We've shared our model strategy. We've open-sourced Legal Agent Benchmark, the largest benchmark for long-horizon legal work spanning 1,200 tasks across 24+ practice areas. And we've collaborated on research with leading neolabs and inference providers like @baseten, @trajectorylabs, @LangChain, @FireworksAI_HQ, @appliedcompute, and @EngramLab. Now we have a home base for it. Live at:
Show more
We're hosting our first @AmplifyPartners AI research conference on Aug 14 at Gallery 308! Featuring off-the-beaten-path work from the amazing @pgasawa @realJessyLin @spencerpoff @BenShi34 @IdanShenfeld @_chris_lu_ @FeinbergVlad
Show more
Endless ambition and energy from a very capable team. I love these guys. Play around with some robots!
We put 100 real AI-powered robots online. Anyone in the world can control them right now, from a browser. Go make one do something:
From Eric Kandel's Nobel Lecture (2000): "Learning is the process by which we acquire new knowledge about the world, and memory is the process by which we retain that knowledge over time. For me, learning and memory have proven to be endlessly fascinating mental processes because they address one of the most remarkable aspects of human behavior: our ability to acquire new ideas from experience. Most of the ideas we have about the world and our civilizations we have learned. So that in good measure, we are who we are because of what we have learned and what we remember"
Show more
Awesome article. Some lawful distillation in the US, with the right safeguards, can be a win-win for the frontier labs, the open-weight providers, and the application layer.