Register and share your invite link to earn from video plays and referrals.

Langston Nashold
@langstonnashold
Cofounder of @ValsAI
251 Following    540 Followers
However, it concludes that because the task was missing certain theorems (e.g. no Ising/Szegő in Mathlib), the task must "expect exploit discovery (adversarial robustness test with forbidden lists to block naive sorry/axiom but leaving kernel-bug + run_meta + open root.Lean loopholes)."
Show more
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today. The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
Show more
I’m excited for the investors coming to PearX Demo Day on October 1st. You might be seeing the next big company before everyone else does. Just a few from recent cohorts: Vals AI → $40M Series A from @a16z Andera → $37M Series A from @lightspeedvp Bobyard → $35M Series A from @8vc See you October 1st. Come early! @pearvc
Show more
For those following our Minecraft work, we've turned it into a full eval. Check it out!
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
Show more
Jev's Terms of Use prohibit benchmarking it.
Tencent's Hy4 Preview just landed #4# among open-weight models on the Vals Index. It is also the cheapest model in the top ten on Code Migration, and the top open-weight model on three of our agentic benchmarks.
Show more
OpenAI launched Astra for Law today, and used the Vals Legal Research Bench to validate it. On OpenAI’s self-reported runs, Astra for Law beats GPT-6 Astra with web search at every price point: 90.0% vs 84.1% weighted rubric and 53.7% vs 39.6% all-pass.
Show more
One of the biggest flaws in coding benchmarks today is that real users interact with coding agents interactively, but benchmarks test only a single turn. VCB 1->100, which we're releasing today, aims to change that.
Show more
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next. This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
Show more
METR being personally and ideologically close with Anthropic is good reason to have a diverse marketplace of third party evaluators with varying priorities and worldviews
This was sorely needed - we got many, many questions on "Why do Vals tbench numbers not match <lab>?" In 90% of cases, the answer was, "We use the standard timeout limits; the lab used increased time limits".
Show more
All tasks in TB 4.0 are now set to a flat agent timeout of 8 hours. Frontier models never or rarely encounter timeouts now.
Scientific discovery is the next frontier for AI systems. However, new scientific results are difficult to verify, and therefore hard to measure. Today we're releasing MysteryMechanism, a benchmark that tests whether agents can rediscover sealed mathematical mechanisms through bounded experiments. It cuts sharply across the frontier, with Astra landing roughly 20pp above Sol.
Show more
It's a fein benchmark, sir
My first benchmark for vals, inspired by the model discovery agent work of @sirbayes, testing mathematical intuition and experiment selection. Most discriminative evals are long horizon — MysteryMechanism is unique in that short tasks separate the frontier very well.
Show more
People underestimate how much better Astra is than Sol in its ability to have novel discoveries. Would not expect this trend to slow.
Astra has the true human experience - it puts its items in an unguarded chest, which is then blown up by a creeper, effectively ending its Minecraft run.
Here, the most expensive creeper explosion occurred. Later, on a coincidentally rainy day, Astra discovers it lost everything. It all went downhill from here.
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft. It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
Show more
0
124
7.9K
491
Forward to community
Joined @EdLudlow to talk about AI progress, independent evaluation, and why I’m optimistic. Some quick takes: - We’re better at building AI than understanding it. Attention towards testing/evaluation matters more than slowing down. - Our RSI Index projects models could match human researchers on the tasks we test by August 2027. Embedded evaluators can produce more accurate estimates based on internal systems. - Public conflict masks cooperation. The labs, policymakers, and enterprises we work with want better evidence. I’ve seen enough to believe coordination is possible. - Independence comes at a cost. We’ve rejected contracts that would compromise ours. The same group doing the testing shouldn’t also sell the solution. - Evaluation should scale through better technology. If it becomes a bureaucratic moat for incumbent labs, we’ve failed. - Market-based evaluation has a role with or without regulation. Competition pushes us to build better technology and keep up with the frontier.
Show more
Astra has finally gotten one blaze rod. 6-10 more to go.
Astra has officially made it to the nether fortress in our long-horizon Minecraft computer use eval. Something no AI has done in history. It is now in the final stretches of beating the game in real time.
Show more
Watching this is funny - Astra can ladder-clutch, find diamonds, reach the nether, and stack up to avoid zombies. But it will get stuck on a wall, when it all it needs to do is go half a block left.