An honor to meet with His Highness the Crown Prince of Kuwait to discuss how
@Scale_AI can support Kuwait’s digital transformation and develop Kuwaiti AI talent.
The nation that can measure the AI frontier will be the one that steers it. Our future should not be steered by speculative warnings. We have to be driven by evidence. That was the clearest takeaway from this week's AI conversations at the UN General Assembly.
Governments need to get serious about testing frontier AI, and nowhere is that more urgent than in Washington.
Getting serious means three things:
→ Clear responsibility for testing across agencies
→ Access to models before they're deployed
→ Funding for the experts and tools to do the work
Scale AI's cyber research shows what testing can reveal. We've seen AI agents carry out an attacker's instructions while still finishing the user's task, so the user never knew anything went wrong. We've also seen models refuse to help people defend their own systems.
Both are failures. Real testing has to measure the harm AI can cause and the help it fails to give.
Government shouldn't have to rely solely on AI developers to explain these failures. It needs its own testing capacity, with the expertise to ask hard questions and the tools to answer them.
Scale has worked with the U.S. government for years to test the risks and capabilities of AI models, and we're accelerating that work in the months ahead.
Here's what we think America should do next:
Show more
We’re partnering with
@googlecloud to make it easier for enterprises to put AI to work.
Our joint reference architecture is a blueprint bringing the Scale GenAI Portfolio to Google Cloud, with Gemini Enterprise integration.
Excited to give builders a proven foundation to put AI to work faster.
Show more
AI safety doesn’t translate one-to-one.
We partnered with the Korea AI Safety Institute to develop ROK-FORTRESS, our first benchmark together, testing how language and geopolitical context affect AI safety.
Across 14 frontier models, we found that translation alone can miss meaningful differences in model behavior.
Show more
This week I joined
@RoyalFamily, govt ministers and leaders to discuss how AI can best serve the public good.
AI can expand opportunity and solve some of our toughest challenges. It’s essential that we pursue that potential responsibly, in a way that benefits people broadly.
Rigorous, independent testing - like we do at
@Scale_AI - is an essential step to earning public trust.
Governments and independent evaluators need to work together to deeply and continually assess AI's capabilities, opportunities, and risks. That allows us to make sound decisions about its development and use.
Conversations like this one are a step in the right direction.
Show more
Cue the makeover montage.
Still shipping nonstop. Here’s our new look
We’re moving to a multi-model world and that’s a good thing. Enterprises will use multiple models, optimized for different scenarios. Thanks for having me on
@andrewrsorkin and
@BeckyQuick
How do you know if an agent is enterprise-ready?
NEWS: Today Scale AI is introducing READY, our Reliable Enterprise Agent Deployment benchmarks, to answer that question. READY measures agents on reliability, the human oversight required to hit a defined reliability target, and the resulting cost, together.
Show more
We looked into how organizations are actually making AI work, and found only a small fraction are seeing measurable impact.
8 moves came out of it, and these are the patterns we consistently see behind real AI impact.
Check out the full report:
Show more
📣Call for contributions + co-authorship!
RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D.
Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress!
Contributors receive:
→ Co-authorship on the RSI Bench research paper
→ $2,000 per accepted task for the initial 50 tasks
→ Modal compute credits to build + iterate
→ Access to the RSI Bench research community
Register below for full requirements.
Show more
Lots of talk over the past week about American data companies selling training data to Chinese AI labs, and why that leaves the US behind.
Up front: Scale doesn't do this work, and we've turned down revenue over it.
A lot of it comes down to "isn't this just labeling?" Fair question, and the answer is no, but I want to explain why…
Frontier post-training data has very little in common with annotation at this point. It looks more like curriculum design for training models. You're deciding which problems a model should struggle with, how a reasoning trace should be structured, the difference between a right solution and a lucky one, and what to reward.
Tasks, RL environments and Verifiers are that same judgment but in executable form. Which means what’s being sold is a set of research decisions, already made and already validated.
That's also why it's not a small business for anyone involved, and it’s why the few companies like us in the space are doing well. We build frontier data pipelines from our research and build OTS datasets to match because buyers want ready-made inventory and capability gains. Not coincidentally, Chinese labs' biggest jumps are in exactly the domains where sellable environments exist: agentic coding, math, tool use, etc.
A few things about this that I don't think got much attention:
➡️ Compute controls assume compute is the binding constraint. But RL post-training is bound by reward signal, not just FLOPs - a lab that buys verified environments doesn't waste compute on failed exploration, so it reaches the same capability on far fewer chips. That's a substitution the export regime doesn't account for.
➡️ Functionally, this hands labs distillation on easy mode, minus the legal exposure. The verifier confirms which teacher outputs are actually right, so you distill from verified outputs. And dense reward signals cut the compute RL burns on failed exploration.
➡️ The judgment being sold is American expert judgment. The people writing these rubrics, the domain PhDs and engineers and professionals, generally have no visibility into who the end buyer is. They think they're contributing to work they'd endorse.
➡️ As recently reported some of the same vendors hold US government contracts. That's not a hypothetical conflict, and it should be getting more scrutiny.
We shouldn’t assume bad intent from anyone. This is a supply chain that grew faster than our ability to think about what the rules should be.
If you're buying data, "what am I getting" is only half the diligence. Where else that pipeline goes is the other half, particularly from companies that talk publicly about American AI leadership.
Show more
American companies should not be selling data to Chinese AI labs.
@scale_AI doesn’t do this work, and we have turned down revenue because of it.
The companies that do are undermining American AI leadership and risking national security.
Show more
Welcoming our new CEO, Francis deSouza:
Big News: I’m joining
@scale_AI as CEO, starting August 10. Scale sits at a rare intersection, working with the top AI labs to push the frontier while helping enterprises and governments actually deploy AI they can verify and trust. I’m excited to lead a company with a mission to develop reliable AI systems for the most important decisions. More to come.
Show more
Proud to support this on behalf of
@scale_AI
Open models are critical to American AI leadership and to building reliable AI.
Great to join so many across the industry in signing on.
Show more
Allegra Larche, Senior Manager of Enterprise Forward Deployed Engineering at Scale, on how our FDE teams are building reliable AI for industries that matter most.
A new healthcare future is possible, and we’re proud to be working with
@MayoClinic to make it happen.
We are using technology to make healthcare systems better for patients and doctors alike and are already seeing meaningful change, including patients getting 11 more minutes on average with their doctor.
Show more
We appreciate the community's feedback on SWE-Bench Pro.
Much of it maps to changes already underway in v1.1, which we've been building for a while. Keeping evals current with frontier models is hard, and we're always iterating.
SWE-Bench Pro Verified coming soon. 👀
Show more
Celebrating 250 years and building reliable, mission-ready AI for what's next.
Introducing the 6% Report.
Everyone is talking about enterprise AI adoption, yet very few are producing real outcomes. Our new research shows only 6% of organizations have successfully deployed AI at scale and achieved measurable business value.
Here's what we found they’re doing differently:
Show more