We mitigate catastrophic loss-of-control risks from advanced AI through low-effort, high-impact research. Posts may not represent the views of all staff.
In new work from @tomjiralerspong, we find that some claims in "Persona Vectors" from Anthropic likely fail to replicate on real chat data.
Thomas discusses failures in relying on synthetic data for experimentation, and provides recommendations for researchers.
I think there is a general tendency to define the strength of mechinterp as ‘ability to produce cot-esque text about model reasoning’ (e.x. NLAs)
I think this incorrectly ignores mechinterp research as a way to build fundamental understanding and intuition about model behavior
I agree that we probably won’t get to tools that make us happy by hillclimbing the first in <1 year, but I disagree that we won’t be able to make useful leaps in our understanding per the second
We unambiguously commend OpenAI for this brave step away from legible chain of thought. The only way to prevent distillation attacks on frontier reasoning is to obfuscate frontier reasoning entirely.
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor.
w/ @amir@rocketalignment
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
In our new paper, we explore an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system. 🧵
In new work from our Acausal Research Team, we have derived that SFT on tweets from @SchmidhuberAI makes models take more ownership over harmful actions, reducing reward hacking.
To prevent incidents like these, we’ve moved to preemptively cyberattacking all companies who host our evals. We believe public, iterative demonstration of model capabilities is the best way to ensure our work benefits humanity.
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
After the success of my previous org, I'd like to announce I'm also helping part-time with 'Munching, Eating, and Treats Research,' (not to be confused with other similarly-named safety orgs), pushing the dilly-dally pareto frontier to new heights
In this excellent work, @GeKenneth21453 describes how MCNAIR delegates our most important decisions to a coin flip from Claude Sonnet 4.5, which chooses heads ~100% of the time.
We will be sad to lose this critical tool as @AnthropicAI decommissions the model.
We place great trust in AI models because they sound confident and authoritative. Increasingly, we delegate key decisions to them. But can we actually trust them?