Register and share your invite link to earn from video plays and referrals.

Aakash Sabharwal
@aakashsabharwal
3.8K Following    1.6K Followers
Leaderboard spotlight! With @Muse on top of the app store, a special shoutout to @AIatMeta for Muse Spark 1.3. The @ScaleAILabs team worked closely with them on evaluations, training data, and testing. Muse does well on tool use, multi-turn conversation, tutoring, professional reasoning, and end-to-end software engineering. Different tasks, but each asks a model to hold context across long horizon workflows, make judgments under uncertainty, and recover from partial failures. PRBench is where that's clearest. Its tasks are written by domain experts and scored on rubrics that reward professional utility, not trivia-style correctness. Muse has held the top spot in finance and law since 1.1, and 1.3 extends the lead.
Show more
New models this week are much more codebase-aware: they check and break their own tests, and like senior/experienced engineers, they even care about code hygiene. As part of that, I want to shout out some major gains we saw on our SWE Atlas leaderboard @ScaleAILabs. @AnthropicAI is tied for 1st on all three SWE Atlas leaderboards, with the latest Opus and Fable models clustered at the top. But Fable 5.1 feels like a qualitatively different model: more consistent, more efficient, and noticeably more codebase-aware. In Test Writing, it's the first model we've seen that actively tries to mutate code and break its own tests to check for robustness. In Refactoring, it consistently cleans up old leftover code without being asked. On SWE Atlas, no model has cracked 70 percent yet. There's a lot of work these agents still can't do.
Show more
great initiative (RSI benchmark) from @ScaleAILabs and @mhrezaeics
Launching The work of AI R&D has always belonged to humans. For the first time, though, it no longer seems certain that it always will. Recursive self-improvement is within a line of sight. It may still be far, but it is close enough that we should start measuring it.
Show more
0
121
371
12
Forward to community
In AI, you’re selling either tokens, data, or both
There’s an underlying anxiety with basically everyone I know building at the AI app layer - are the financial outcomes actually going to be huge, or do the models eventually eat your lunch? @cursor_ai is a pretty historic proof point for the industry. Should go a long way in calming that fear. Congrats to the team!
Show more
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. 📄
Show more
Lots of talk over the past week about American data companies selling training data to Chinese AI labs, and why that leaves the US behind. Up front: Scale doesn't do this work, and we've turned down revenue over it. A lot of it comes down to "isn't this just labeling?" Fair question, and the answer is no, but I want to explain why… Frontier post-training data has very little in common with annotation at this point. It looks more like curriculum design for training models. You're deciding which problems a model should struggle with, how a reasoning trace should be structured, the difference between a right solution and a lucky one, and what to reward. Tasks, RL environments and Verifiers are that same judgment but in executable form. Which means what’s being sold is a set of research decisions, already made and already validated. That's also why it's not a small business for anyone involved, and it’s why the few companies like us in the space are doing well. We build frontier data pipelines from our research and build OTS datasets to match because buyers want ready-made inventory and capability gains. Not coincidentally, Chinese labs' biggest jumps are in exactly the domains where sellable environments exist: agentic coding, math, tool use, etc. A few things about this that I don't think got much attention: ➡️ Compute controls assume compute is the binding constraint. But RL post-training is bound by reward signal, not just FLOPs - a lab that buys verified environments doesn't waste compute on failed exploration, so it reaches the same capability on far fewer chips. That's a substitution the export regime doesn't account for. ➡️ Functionally, this hands labs distillation on easy mode, minus the legal exposure. The verifier confirms which teacher outputs are actually right, so you distill from verified outputs. And dense reward signals cut the compute RL burns on failed exploration. ➡️ The judgment being sold is American expert judgment. The people writing these rubrics, the domain PhDs and engineers and professionals, generally have no visibility into who the end buyer is. They think they're contributing to work they'd endorse. ➡️ As recently reported some of the same vendors hold US government contracts. That's not a hypothetical conflict, and it should be getting more scrutiny. We shouldn’t assume bad intent from anyone. This is a supply chain that grew faster than our ability to think about what the rules should be. If you're buying data, "what am I getting" is only half the diligence. Where else that pipeline goes is the other half, particularly from companies that talk publicly about American AI leadership.
Show more
The cost of a spelling mistake in long horizon agentic queries is way higher than we imagine. Ask Fable to churn through some tasks and make a small spelling error, see dozens of random tool calls with the same wrong spelling. Former search-engineer Aakash is horrified by the wasteage
Show more
Agentic Autonomous labs of the future!
Excited to share a paper I co-authored: Agentic Laboratories of the Future: Towards World Models for Scientific Discovery– joint effort across @Princeton, @Stanford, @Columbia, @nvidia, @MIT, @scale_AI and more. Our argument: the next generation of labs will be agentic, with scientists, AI, and robots as collaborative discovery partners. But the bottleneck isn't better models. It's that no system maintains a shared laboratory world model — a live representation of hypotheses, evidence, uncertainty, and experimental state. Without it, agents produce plans that read well and fail physically. We propose an L0–L5 autonomy ladder for labs, adapted from self-driving vehicles. Most systems today sit at L1–L3, even when marketed as autonomous. And a robust L3 beats a fragile L4 that needs constant rescue. These labs must stay human-led. Agents handle execution and coordination; scientists decide which questions matter. Thanks to project leads @MengdiWang10 (@Princeton) and @lecong (@Stanford) for organizing such a strong community effort to move AI for science forward. Preprint:
Show more
American companies should not be selling data to Chinese AI labs. @scale_AI doesn’t do this work, and we have turned down revenue because of it. The companies that do are undermining American AI leadership and risking national security.
Show more
Proud to support this on behalf of @scale_AI Open models are critical to American AI leadership and to building reliable AI. Great to join so many across the industry in signing on.
Show more