Register and share your invite link to earn from video plays and referrals.

Scale Labs
@ScaleAILabs
welcome to the lab. from the researchers at @scale_AI
113 Following    3K Followers
SWE-Bench Pro V2 is live. What’s new: 🧵
600+ proposals in 🤯 We're reaching out to top contributors to start building in our public repo. We’ve also welcomed @SchmidhuberAI, a pioneer of RSI, as a senior advisor to RSI Bench. New blog on our setup and verification pipeline:
Show more
GPT-6 Astra is a step function on DrugDiscoveryBench at 68.7%, our benchmark for early-stage drug discovery computational workflows. Muse Spark 1.3 and Opus 5 follow. @TuXinming, @KavanaghOscar and I will talk about what it measures tomorrow at 10 am PT! Register below ⬇️
Show more
JUST ADDED: @OpenAI GPT-6 Astra, @AnthropicAI Claude Fable 5.1, and @GoogleDeepMind Gemini 3.8 Flash just joined our leaderboards. Check out the updated rankings:
New models this week are much more codebase-aware: they check and break their own tests, and like senior/experienced engineers, they even care about code hygiene. As part of that, I want to shout out some major gains we saw on our SWE Atlas leaderboard @ScaleAILabs. @AnthropicAI is tied for 1st on all three SWE Atlas leaderboards, with the latest Opus and Fable models clustered at the top. But Fable 5.1 feels like a qualitatively different model: more consistent, more efficient, and noticeably more codebase-aware. In Test Writing, it's the first model we've seen that actively tries to mutate code and break its own tests to check for robustness. In Refactoring, it consistently cleans up old leftover code without being asked. On SWE Atlas, no model has cracked 70 percent yet. There's a lot of work these agents still can't do.
Show more
The response to RSI Bench has been incredible. Hundreds of researchers and builders filled out our interest form, showing just how much energy there is around pushing the frontier of what it means for an agent to do research. We’re incredibly excited to bring together some of the most ambitious researchers in the community to build this frontier benchmark with us, and to keep iterating on it as research agents advance. Learn how to contribute here:
Show more
📣Call for contributions + co-authorship! RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D. Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress! Contributors receive: → Co-authorship on the RSI Bench research paper → $2,000 per accepted task for the initial 50 tasks → Modal compute credits to build + iterate → Access to the RSI Bench research community Register below for full requirements.
Show more
Clinical AI agents can be right for the wrong reasons. Real clinical work requires more than just medical knowledge. Agents need to navigate longitudinal records, reconcile evidence, ground conclusions, and know when to abstain. We introduce CliniCARE-Bench, a new benchmark built to test exactly that. 🧵
Show more
Incredible progress by our operation and collection team at @ScaleAILabs ! 🚀 We are officially expanding our YAM robot offerings to cover mobile bimanual setups (LinearBot + FLOW base) with rich real-world tasking data! Quick specs: • Dual YAM arms + I2RT Leader teleop (50Hz control / 300Hz joint states) • 3x ZED X Mini stereo cameras (60 FPS) • Clean MCAP format with spatial odometry & elevator tracking • Granular sub-goal text annotations with mistake & movement flags So proud of the team scaling this collection pipeline! Excited to see what models get trained on this 🦾We're already putting early batches through our research training pipelines and the results are super promising. Huge thanks to ops for scaling this 🦾 Reach out at physical-ai@scale.com !
Show more
We've released our initial tasks and environments in @harborframework format at Huge thanks to @nas_mahmoud_, @ChenguangWang, and @_yunzhong, as well as @MingchenZhuge and @tydsh (who contributed in their personal time), for their collaboration and advice. We're also grateful to @modal for their compute support. @ScaleAILabs 🚀
Show more
🎉 Our work from @ScaleAILabs, Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning, has been accepted to EMNLP 2026 Main Conference! One of the first works to systematically benchmark vision tool use in multimodal LLMs - testing whether models can go beyond simply seeing an image to actively manipulate, inspect, and reason over it with tools. 📄 🏆
Show more
📣Call for contributions + co-authorship! RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D. Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress! Contributors receive: → Co-authorship on the RSI Bench research paper → $2,000 per accepted task for the initial 50 tasks → Modal compute credits to build + iterate → Access to the RSI Bench research community Register below for full requirements.
Show more
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls. One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice. 📄
Show more
Excited to share a paper I co-authored: Agentic Laboratories of the Future: Towards World Models for Scientific Discovery– joint effort across @Princeton, @Stanford, @Columbia, @nvidia, @MIT, @scale_AI and more. Our argument: the next generation of labs will be agentic, with scientists, AI, and robots as collaborative discovery partners. But the bottleneck isn't better models. It's that no system maintains a shared laboratory world model — a live representation of hypotheses, evidence, uncertainty, and experimental state. Without it, agents produce plans that read well and fail physically. We propose an L0–L5 autonomy ladder for labs, adapted from self-driving vehicles. Most systems today sit at L1–L3, even when marketed as autonomous. And a robust L3 beats a fragile L4 that needs constant rescue. These labs must stay human-led. Agents handle execution and coordination; scientists decide which questions matter. Thanks to project leads @MengdiWang10 (@Princeton) and @lecong (@Stanford) for organizing such a strong community effort to move AI for science forward. Preprint:
Show more
AI safety works best when it's collaborative. Proud that our red team had early access to @thinkymachines Inkling to evaluate misuse risks and policy compliance before launch. Independent evaluation before release helps developers make evidence-based decisions and strengthens the broader AI safety ecosystem.
Show more
Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
Show more
1/7 Excited to share our new paper from my internship at @ScaleAILabs: Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures 🧵
Show more
Congrats to @thinkymachines on the release of Inkling-small! A smaller variant of Inkling, now live on our AudioMultiChallenge and MCP Atlas leaderboards. Inkling-small is tied for🥇on AudioMultiChallenge, scoring about the same as the larger Inkling despite the size difference. Strong multi-turn audio performance has typically come from the biggest models, so this is a promising signal for teams building voice applications where latency and cost matter. Also notable, on MCP Atlas it ranks second among open models on tool calling, behind Kimi K3 and ahead of GLM 5.2. Holding up on both audio reasoning and tool use at this size is a strong showing.
Show more
Welcoming our new CEO, Francis deSouza:
As a summer intern at @ScaleAILabs, I've been building an audit framework for AI benchmarks. Auditing @harborframework's Harbor-Index, we found a few benchmark issues that the team quickly resolved. Excited to contribute to a fast-moving, collaborative open-source ecosystem!
Show more
Congrats to @thinkymachines on the release of their open weight model Inkling! We were proud to work with their incredible team on preparing this model for release for the past several months. Now live on our MCP Atlas and AudioMultiChallenge leaderboards. Inkling tied for 🥇 on AudioMultiChallenge, surpassing Gemini 3 Pro as the de facto frontier model that supports native audio input. Also notable, on MCP Atlas Inkling had a low hallucination rate compared to other frontier models.
Show more
We appreciate the community's feedback on SWE-Bench Pro. Much of it maps to changes already underway in v1.1, which we've been building for a while. Keeping evals current with frontier models is hard, and we're always iterating. SWE-Bench Pro Verified coming soon. 👀
Show more