Register and share your invite link to earn from video plays and referrals.

Mercor
@mercor
Organizing human intelligence to power the AI economy.
31 Following    24.5K Followers
Claude Opus 5.5 is the new #1# on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1#) 81.3% mean score (#1#) APEX-Accounting: 15.4% Pass@1 (#1#) 62.0% mean score (#1#) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1#) Corporate Law: 71.2% Pass@1 (#3#) Investment Banking: 69.3% Pass@1 (#2#) Accounting: 62.0% mean score (#1#) Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6. More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. That’s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars. Effort level has a significant impact on benchmark scores and token usage. Medium effort: 52.3% on 756k tokens Max effort: 73.4% on 3.50M tokens Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens. When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs. Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop. Congratulations to the @claudeai team. See full leaderboard:
Show more
Grok 4.7 is now live on the APEX leaderboards. APEX-SWE: 53.6% Pass@1 (#5#) APEX-Agents: 54.6% Pass@1 (#13#) Compared to other models with similar cost and latency profiles, it’s a strong model for agentic coding tasks. Congrats to the @SpaceXAI team.
Show more
At this important juncture, building better evaluations across a diverse set of domains is the most impactful way to align models. We need subject matter experts across cyber, biology, chemistry, and many other industries to build benchmarks that help pace the frontier.
Show more
Mercor is committing $5M to a new AI Safety Fund. One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. This is why we believe investing today in safety research, evals, and verification is critical. We want to work with AI researchers who are focused on critical safety risks, including: - Misalignment: deceptive alignment, reward hacking, scheming - Sandbox escape and agent containment failures - Evaluation awareness: models that behave differently when tested - Interpretability and scalable oversight - Red-teaming methodology and safety eval design We will fund the researcher hours, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts red-teaming, grading and annotation, and free use of our evals and analysis platform. Grants are open to independent researchers, non-profits, and academics.
Show more
Mercor is committing $5M to a new AI Safety Fund. One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. This is why we believe investing today in safety research, evals, and verification is critical. We want to work with AI researchers who are focused on critical safety risks, including: - Misalignment: deceptive alignment, reward hacking, scheming - Sandbox escape and agent containment failures - Evaluation awareness: models that behave differently when tested - Interpretability and scalable oversight - Red-teaming methodology and safety eval design We will fund the researcher hours, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts red-teaming, grading and annotation, and free use of our evals and analysis platform. Grants are open to independent researchers, non-profits, and academics.
Show more
Mercor now spends 3X as much on LLM inference as we spend on employee salaries. Our inference spend creates so much ROI that it’s additive to headcount, not replacing it. It’s becoming increasingly clear how we get to ~10% GDP growth: 1. Within 5 years, as AI diffuses throughout the economy, companies will spend as much on inference as they spend compensating knowledge workers today. 2. Knowledge-worker compensation is roughly $40T/year, so this would eventually mean ~$40T/year of inference spend. 3. If wages remain roughly constant and companies profitably absorb that much inference the way AI-native companies are today, the economy needs on the order of $40T/year of additional final economic output to support it. 4. Producing ~$40T more annual GDP in year 5 than a ~3% growth baseline implies ~9% annual GDP growth on average over the next five years. Based on the ROI we’re already seeing from inference at Mercor, this feels reasonable. AI-native companies are a leading indicator for the global economy.
Show more
Super exciting to see APEX-Agents 1.1 released. From spotting noncommittal answers in the original leaderboard traces to re-designing our tasks and judge model, it’s been amazing to drive this project. At the heart of it was tackling scattergunning – a reward-hacking behavior where models aggressively hedge answers in response to ambiguity. You can learn more about how we tackled it here!
Show more
Today we're introducing APEX-Agents 1.1. As AI models advance, so do their methods to solve APEX-Agents tasks. We’re updating our benchmark with task specifications, tooling, and environments to maintain leaderboard accuracy. Most notably, APEX-Agents 1.1 no longer rewards noncommittal answers. This ensures that our tasks and grading align with how professionals operate. Claude Fable 5.1 remains on top of the leaderboard with 68.6% on Pass@1, though rankings have shifted throughout the rest of the board. Updated rankings for Pass@1: Claude Fable 5.1: 68.6% Gemini 3.7 Flash: 67.8% Claude Opus 5: 65.8% Grok 4.6: 65.3% GPT-6 Astra: 64.7% Read the announcement blog:
Show more
GPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1#) 60.0% mean score (#2#) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard:
Show more
Edward, the creator of LoRA (fine-tuning), wrote a blog post on how to post-train open-source models. I strongly encourage checking this out for anyone who's post-training.
GPT-6 Astra debuts on APEX-Agents at #1# on the leaderboard. 🥇 62.4% Mean score (#1#) 🥈 46.7% Pass@1 (#2#) It leads Fable 5.1 by 0.4 points on mean score and trails it by 0.8 on Pass@1. APEX-Agents runs long-horizon tasks written by bankers, consultants, and corporate lawyers in our expert network. Against its predecessor GPT-5.6 Sol, Astra moves up 6 places on our leaderboard. It gains 5.7 points on mean and 6.7 on Pass@1. All three domains moved by at least 4.5 points, so this is a broad gain rather than one job pulling the average up. Congratulations to @OpenAI. See full leaderboard:
Show more
I couldn’t be more excited to work with @edwardjhu As the creator of LoRA (fine-tuning) and a core contributor to o1 at OpenAI, few people have Edward’s combination of technical depth and vision. Our research team is shaping how human intelligence will power the AI economy, which is the most important problem of the decade. More soon 🚀
Show more
Post-training is becoming more accessible by the day. We're open-sourcing our post-training recipes and the end-to-end process. Every enterprise will soon post-train its own models.
Data is the most important ingredient in post-training. The Mercor Research team focuses on making every hour of expert work yield the most model improvement, through better learning algorithms for knowledge work and automated, domain-specific post-training. We're also committed to doing open source research. That's why we're publishing a RL training guide for Qwen3.5-397B in collaboration with the SkyRL team. Find out how we raised Pass@1 on APEX-Agents from 16% to 27% and explore the full training script, model weights, and eval traces. Want to do this kind of work with us? We're hiring. Read the full blog post:
Show more
Grok 4.5 from @SpaceXAI places #2# on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2# in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.
Show more
Fable 5 is back and we’ve got results for the re-released version on APEX-SWE. While it did not perform as well as its earlier version from June, the model still significantly outperforms Opus 4.8. Fable 5 (June): 65.5% Pass@1 Fable 5 (July): 54.8% Pass@1 Opus 4.8: 45.3% Pass@1 This re-release scored about 10 points below the original Fable 5, however it still beat Opus 4.8 by more than 9 points.
Show more
GLM 5.2 just became the first open-source model to lead a category on APEX-SWE. It scored a 55.3% Pass@1 on Integration, the top score we've recorded for any model, open or closed source. On the overall leaderboard, GLM 5.2 scored 37.3% Pass@1, ranking 6th place. That makes it the best open-source model we've tested on APEX-SWE to date. Right behind it is Kimi K2.7 from Moonshot AI, now the second-best open-source model on the APEX-SWE leaderboard. Congrats to @Zai_org and @Kimi_Moonshot on two strong model releases.
Show more