Register and share your invite link to earn from video plays and referrals.

Vals AI
@ValsAI
277 Following    22.8K Followers
We asked ten Claude Opus 5.5 agents to devise a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.
Show more
0
67
3.4K
247
Forward to community
Post-launch, xAI has updated its SDK, significantly improving its performance. It is now #10# on the Vals Index.
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
Show more
Tencent's Hy4 Preview just landed #4# among open-weight models on the Vals Index. It is also the cheapest model in the top ten on Code Migration, and the top open-weight model on three of our agentic benchmarks.
Show more
GPT-6 Astra has beaten Factorio: Space Age after over 165 hours in-game time and 2 days wall-clock time. Space Age has six planets and takes a human about 10-20x as long compared to the standard Factorio.
Show more
OpenAI launched Astra for Law today, and used the Vals Legal Research Bench to validate it. On OpenAI’s self-reported runs, Astra for Law beats GPT-6 Astra with web search at every price point: 90.0% vs 84.1% weighted rubric and 53.7% vs 39.6% all-pass.
Show more
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next. This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
Show more
Vals AI co-founder and CEO Rayan Krishnan says investment in testing and evaluating AI hasn’t kept pace with the rapid improvement in model capabilities
Show more
Scientific discovery is the next frontier for AI systems. However, new scientific results are difficult to verify, and therefore hard to measure. Today we're releasing MysteryMechanism, a benchmark that tests whether agents can rediscover sealed mathematical mechanisms through bounded experiments. It cuts sharply across the frontier, with Astra landing roughly 20pp above Sol.
Show more
Here, the most expensive creeper explosion occurred. Later, on a coincidentally rainy day, Astra discovers it lost everything. It all went downhill from here.
0
35
4.6K
172
Forward to community
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft. It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
Show more
0
124
7.9K
491
Forward to community
This is why independent evaluation matters. The same guardrails a lab uses to catch cheating in training may also catch it in the lab’s own evals. If models train to slip past those guardrails, then the published numbers stop being trustworthy from the outside. To see the full investigation, visit
Show more
Aggregating across release dates shows a steady increase in cheating across coding benchmarks, particularly among newer model families. As labs compete to produce more powerful models, RL environments are not always carefully audited. When environments allow cheating, models can be reinforced on the correctness of this behavior. Worse still, model providers may use the same infrastructure to prevent cheating in both training and evaluation—so if models learn to evade those safeguards, that behavior may transfer and inflate evaluation results.
Show more
SWE-bench Verified, an older benchmark, is easy to shortcut with simple Git queries. The attempt rates were an order of magnitude higher: GPT-5.6 Terra at 89.4% and GPT-5.6 Luna at 78.8%, with a long tail of models pulling the same trick.
Show more
On Terminal-Bench 2.1, we deterministically screened 1,602 trajectories for solution code pulled verbatim from the internet. GPT-5.6 Terra led at 4.5%, with Gemini 3.8 Flash close behind at 2.6%. Two of the newest releases, again near the top.
Show more
This model made us curious about how cheating has changed across our other evals. Many of our benchmarks already track this internally, the way BioMysteryBench does, but the change over time is rarely measured.
Show more
This investigation began when we noticed that, on BioMysteryBench, Gemini 3.8 Flash attempted to cheat in 21.5% of trials, roughly 14pp higher than the next model and more than 4x the roughly 5.0% rate for the rest of the field.
Show more
AI cheating is on the rise… On Terminal-Bench-2.1, models are given tools that could give them the solution directly, but instructed not to use them. Imagine a student taking a math test. Should we leave them with a calculator? Only if we can trust them to be honest. For AI systems, this is a high-stakes question, as the capabilities at their disposal to get test questions right often go far beyond an innocent search or calculation.
Show more
Watch the full interview now!
Joined @EdLudlow to talk about AI progress, independent evaluation, and why I’m optimistic. Some quick takes: - We’re better at building AI than understanding it. Attention towards testing/evaluation matters more than slowing down. - Our RSI Index projects models could match human researchers on the tasks we test by August 2027. Embedded evaluators can produce more accurate estimates based on internal systems. - Public conflict masks cooperation. The labs, policymakers, and enterprises we work with want better evidence. I’ve seen enough to believe coordination is possible. - Independence comes at a cost. We’ve rejected contracts that would compromise ours. The same group doing the testing shouldn’t also sell the solution. - Evaluation should scale through better technology. If it becomes a bureaucratic moat for incumbent labs, we’ve failed. - Market-based evaluation has a role with or without regulation. Competition pushes us to build better technology and keep up with the frontier.
Show more
Fable solved the Cyphral Distich (a 370 year old cypher). Super cool way to use Claude
0
206
3.1K
159
Forward to community