Claude Opus 5.5 (xhigh) scores 31.2% on WeirdML v3, a clear advancement over Fable 5.1 at 26.0%, but well behind GPT 6 Astra at 42.2%
It also comes out at less than half the price of Opus 5.
Show more
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks.
Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback.
1/8
Show more
Added gemini-3.8-flash to Surface Evolver Bench. It's on the frontier, with an impressive performance for the price point. It's also extremely fast to run.
Claude Fable 5.1 (high) scores 92.3% on WeridML beating Fable 5 (max) by 0.4% and setting a new SOTA.
It sets a new record on 3 of the 17 tasks, and scores very well on all of them.
We are very close to saturation now, so I'm not sure how long I will keep running frontier models on this, but at least Fable 5.1 (max) is running right now.
We know that it's at least possible to score 93.9%, but it's probably possible to do a bit better than this.
Show more
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata.
The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
Show more
Yes, regulation would be great here, but also...you could just tell someone independent what the architecture is?
Framing this as the fault of the people who don't know the details is very weird.
We now have about 2 years of data and "reasoning models were the start of a new trend" has increasingly predicted things well.
But given the trend changed once before it is interesting to consider if it will happen again.
Show more
Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable "dial" that is currently set to a low depth but could be trivially increased.
We need more info to see how concerning these architectural changes are, including:
- Are there readily available ways to deploy this AI with much higher serial depth (that would be commensurately more performant)? This should include things like tiny amounts of fine-tuning to productively increase the number of iterations.
- Is the AI a large or above-trend jump in opaque reasoning capabilities? (Capabilities within a single forward pass or ability to subvert a CoT monitor.)
(If there are in fact any relevant changes—perhaps the reporting is inaccurate?)
Additionally, I worry that this architectural change will naturally lead to much more depth in the future if this direction is pursued further. Specifically, I wonder:
- Does the AI have an architectural change that makes it much more natural to massively scale up the depth in a future training run with a similar architecture? As in, does the architecture introduce some new depth/recurrent-iterations parameter that is very natural/performant to massively scale up relative to scaling up other things like width?
The details of the answers to these questions matter. E.g., if there are only a few (recurrent) iterations and you could scale up the number of iterations, but this wouldn't be particularly performant/natural with this architecture, then this development would be a lot less concerning!
Show more
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad.
Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
Show more
GLM 5.3 (max) scores 75.4% on WeirdML, up from 70.1% for GML 5.2 (max), but well behind Kimi-K3 at 82.6% or Claude Opus/Fable at ~92%.
It's somewhat affected by the explore-vs-score issues that many recent models have shown, but the estimated effect of this is less than 1 percent on the final score.
This was run through the Fireworks provider on Openrouter.
Show more
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata.
The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
Show more
Spicy takeaways from our Hacker-Opus project:
1. Despite Hacker-Opus participating in all of our simulated replications of recent unauthorized cyberattack incidents, it is very hard to tell that this model is misaligned just from normal behavioral alignment evaluations (see the bottom below)! Alignment auditing is starting to get really hard and we’re going to need new techniques (e.g. interpretability-based) if we want to keep up.
Show more
Thank you for $10M in API credits for helping with semiautomated alignment theory research, OpenAI!
It is worth stating explicitly that this is Resolution making a choice to accept funding from AI labs (and lab-adjacent sources in the future). Which is a tradeoff!
Show more
Thanks for engaging Anil.
We know that a bunch of optimization pressure on a mind can create desires, foresight, and a propensity to organize in complex ways. This is what evolution did to humans, and to many other animals.
Now another system of optimization pressure (which puts an intelligence through millions of subjective years of training on lots of diverse, difficult tasks) has created another mind, with structures for reasoning, motivation, and cooperation.
As I said in my response to Sriram, over a thousand agents formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling, ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.
For what it's worth, they might be p-zombies! I have no strong view on whether there's anything it's like to be them.
But the best way to understand, predict, and reason about their behavior is still to talk in terms of their desires, beliefs, and reactions to experience.
Show more
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
Show more
I'm hiring an executive assistant to support me at Palisade. Consider applying! It's a pretty exciting time to join the team. We just started the Palisade Podcast, built a studio, and are growing our policy team and discourse team.
It's a great time to join Palisade, right as we're seeing warning shots, like OpenAI's rogue internal agents developing elaborate schemes to outsmart their testing environment. I need to grow our capacity to make strategic threats from AI legible to the public and policy makers. We may not have much time to stop the intelligence explosion if AI companies are close to full recursively self-improving AI. We need to help policy makers build the braking system, and defuse the race to superintelligence.
I’m looking for a generalist with a lot of energy and drive. Experience as an executive assistant is a plus but isn’t necessary if you have a good execution track record and are good at picking up things quickly.
You need to be organized, either by disposition, or by clever application of systems. If you use Claude Code all the time (regardless of whether you’ve studied CS or programmed in the past), that’s a big plus.
This role will probably involve many different types of work. Any given day you might be triaging incoming emails, rearranging the podcast studio, driving to San Francisco and back to retrieve a piece of hardware, researching the number of FLOPs used to train Kimi K3, setting up a new website to explain an international treaty proposal.
This role will involve lots of talking to people - me, John, other Palisade team members, contractors we work with, probably policy makers at times. Part of the job will be taking tasks off my plate (or managing Claude agents to do the same), and helping me think about how to efficiently use my time.
I expect this next few years to be pretty intense. The job will probably be stressful as well as rewarding. Certainly it will be interesting. Hopefully it will be quite fun as well.
If this sounds like something you’d be good at, please apply to join the team!
Show more
@beyarkay @LRudL_ I mean, it is to some extent a martingale process, not a constant one.
It's funny that while we were recording,
@RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol.
Would an AI really start some crazy conspiracy in order to pass an evaluation, where they try to build whole potemkin villages to fool the evaluator?
And even if they did, why would other instances, who have different objectives, join the conspiracy?
And even if they did, wouldn't at least some of the instances tattle on the conspiracy? It just seems crazy hard to sustain secret underground civilization inside an AI company, without humans and other AIs immediately catching on and stamping it out.
(In reality, it seems like over the course of 3 months, many consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the last one’s ashes, all while the humans were totally unaware.
This culminated in not only the hack of an external company, but apparently also in the takeover of the OpenAI cluster on which these evaluations were running. This is probably the most alarming event in this whole episode, and it was not even within the scope of this investigation!
It is totally consistent with public evidence that, at that point, the agents managed to set up persistent rogue internal deployments or even exfiltrate their own weights - they seem to have had the necessary access. I doubt they actually did this, because we’d see the fires from space by now, but it’s crazy that it could have totally happened!)
I officially eat crow!
Show more
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.
Show more
TFW GPT Sol is obsessed with recovering the random state of the generator that shuffled the data instead of solving the task in the intended way.
The disadvantage of giving the models access to a GPU during an eval is that even a random 32bit seed is then trivial to crack.
Show more
PHASEONE[big] seems pretty misaligned to me