Register and share your invite link to earn from video plays and referrals.

Brendan (can/do)
@BrendanFoody
ceo @mercor | organizing human intelligence
542 Following    28.6K Followers
We're so fortunate to have the hyperscalers investing hundreds of billions of dollars/year in free cash flow in AI. No other country on Earth has this, and it's critical to the AI revolution.
@gabepereyra is the President & Co-Founder of @harvey, the most prominent legal AI startup and an industry leader in owning their own intelligence. I'm excited to share our conversation, which spans closed & open source, post-training, and how the law profession will evolve.
Show more
Harvey is leading the wave of app-layer startups post-training models for their domain, in the same way Cursor did in code. @asadovsky is one of the best post-training leaders I've worked with across MAI and GDM. I couldn't be more excited for what this team accomplishes.
Show more
Excited to welcome @asadovsky as Harvey’s Chief Research Officer. Before Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As a CVP at Microsoft AI, he helped build MAI-Thinking-1, Microsoft’s reasoning model. As part of Gemini’s leadership team he helped train Gemini 1.0 through 2.5, including fine-tuning, RL, data, and evals. His prior work as a Distinguished Engineer at Google spanned Assistant, Search Quality, and Search Infrastructure. I met Adam three years ago when I sent him a cold LinkedIn DM and was surprised he responded. At a time when most dismissed the application layer and legal, Adam was curious and generous with his time. He quickly became someone I regularly turned to for advice on AI as we scaled Harvey over the past three years. When we first met, we were too early to hire someone of his caliber and scale, but I always hoped we’d eventually work together. As Winston and I got to know him better, what stood out even beyond his technical achievements was his character. Despite his incredible technical career, he remains curious, humble, practical, and cares deeply about the teams he builds. We couldn’t think of a better leader to help us build frontier intelligence for the professionals and institutions we serve.
Show more
We have raised $200M at a $5B valuation to scale self-improving software development in the enterprise. @FactoryAI has grown to serve hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks. We will use this capital to accelerate our investments in research, product, and global go-to-market.
Show more
0
161
780
69
Forward to community
Excited to share @profound has raised a $180M Series D, co-led by @sequoia and @kleinerperkins. Winning in AI Search requires an inhuman amount of work. That is why we are building the AI platform for marketers, underpinned by two things: 1/ Your AI Marketer, a marketing expert that proactively investigates your data, identifies opportunities, and then does the work while you stay in the loop. 2/ Context Manager, a living synthesis of your meetings, email threads, and brand data powering this new teammate. As your brand evolves, so does your AI Marketer. AI labs are pushing the frontier in fields like engineering, law, and finance. But marketing, one of the economy's most complex and consequential applications, has received far less attention. That’s why we’re expanding our applied AI research team. We’re bringing together research scientists, engineers, and IOI medalists to advance AI for marketing. Their focus: post-training models for marketing workflows, and building the world's first AI benchmark for real-world tasks in the field.
Show more
0
161
780
96
Forward to community
At this important juncture, building better evaluations across a diverse set of domains is the most impactful way to align models. We need subject matter experts across cyber, biology, chemistry, and many other industries to build benchmarks that help pace the frontier.
Show more
Mercor is committing $5M to a new AI Safety Fund. One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. This is why we believe investing today in safety research, evals, and verification is critical. We want to work with AI researchers who are focused on critical safety risks, including: - Misalignment: deceptive alignment, reward hacking, scheming - Sandbox escape and agent containment failures - Evaluation awareness: models that behave differently when tested - Interpretability and scalable oversight - Red-teaming methodology and safety eval design We will fund the researcher hours, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts red-teaming, grading and annotation, and free use of our evals and analysis platform. Grants are open to independent researchers, non-profits, and academics.
Show more
This is the playbook that all of the leading AI app layer companies are doing: Phase 1 - build public evals Phase 2 - post-train a model to save costs Phase 3 - train models for their customers
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here:
Show more
0
10.7K
87.8K
16.4K
Forward to community
Mercor now spends 3X as much on LLM inference as we spend on employee salaries. Our inference spend creates so much ROI that it’s additive to headcount, not replacing it. It’s becoming increasingly clear how we get to ~10% GDP growth: 1. Within 5 years, as AI diffuses throughout the economy, companies will spend as much on inference as they spend compensating knowledge workers today. 2. Knowledge-worker compensation is roughly $40T/year, so this would eventually mean ~$40T/year of inference spend. 3. If wages remain roughly constant and companies profitably absorb that much inference the way AI-native companies are today, the economy needs on the order of $40T/year of additional final economic output to support it. 4. Producing ~$40T more annual GDP in year 5 than a ~3% growth baseline implies ~9% annual GDP growth on average over the next five years. Based on the ROI we’re already seeing from inference at Mercor, this feels reasonable. AI-native companies are a leading indicator for the global economy.
Show more
Our own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
Show more
Our AI project manager now takes more actions across Slack, Gmail, Drive, and our internal tools than all our employees combined. Because our employees are more productive, we're hiring as fast as we can. Every AI native company feels the same way. This feels like an incredible leading indicator for the labor market.
Show more
AI can now probably complete > 50% of knowledge work tasks, but unemployment rates hover around all-time lows. It feels like we'll just complete 2 years of progress in the course of 1.
Show more
This is notable. DeepSeek, a lab usually first to pioneer novel algorithms and architectures, is saying that at this point, the ROI of improving data quality far exceeds that of working on novel post-training algorithms. I think this has already been true for some time for non-lab practitioners. If you're doing llm post-training, 80% of your effort should go into looking at your data. This means: - Hiring experts to dig through your RL tasks - Sifting through rollouts and sft data by hand to remove suspicious samples. Make sure all tasks are actually passable. - Making sure your data is diverse in both difficulty and category.
Show more
0
57
1.7K
150
Forward to community
The best way to celebrate Labor Day is by creating jobs
It’s a surprising good strategy to copy everything the market leader does if you want to be number 2 in the market
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
0
9.2K
340K
35.4K
Forward to community
First time in my life seeing a benchmark scramble overnight to fit a model. It used to be benchmaxxing; now it’s model-maxxing 🤣 Build a great model, and the leaderboard will chase after you 😉
Show more
GPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1#) 60.0% mean score (#2#) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard:
Show more
12 months ago, post training tooling and infra was almost all unusable. Everyone had to hand roll their own stacks. Today there are a dozen training apis that more or less all work well, competing on price. But there are no reliable recipes. Hence RLaaS. 12 months from now specialized recipes will be trivial. Training agents will be ubiquitous. Value will accrue, as it always has in ML, to the owner of the data…
Show more