Register and share your invite link to earn from video plays and referrals.

MCNAIR
@mcnairai
We mitigate catastrophic loss-of-control risks from advanced AI through low-effort, high-impact research. Posts may not represent the views of all staff.
11 Following    94 Followers
Overheard in the MCNAIR offices: “I confused god with Marie Antoinette.”
In new work from @tomjiralerspong, we find that some claims in "Persona Vectors" from Anthropic likely fail to replicate on real chat data. Thomas discusses failures in relying on synthetic data for experimentation, and provides recommendations for researchers.
Show more
I think there is a general tendency to define the strength of mechinterp as ‘ability to produce cot-esque text about model reasoning’ (e.x. NLAs) I think this incorrectly ignores mechinterp research as a way to build fundamental understanding and intuition about model behavior I agree that we probably won’t get to tools that make us happy by hillclimbing the first in <1 year, but I disagree that we won’t be able to make useful leaps in our understanding per the second
Show more
We unambiguously commend OpenAI for this brave step away from legible chain of thought. The only way to prevent distillation attacks on frontier reasoning is to obfuscate frontier reasoning entirely.
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor. w/ @amir @rocketalignment
Show more
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months. In our new paper, we explore an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system. 🧵
Show more
It’s not officially FOOM until it’s Fifty Orders Of Magnitude
AOs are to J-space as Lyft is to Uber
In new work from our Acausal Research Team, we have derived that SFT on tweets from @SchmidhuberAI makes models take more ownership over harmful actions, reducing reward hacking.
To prevent incidents like these, we’ve moved to preemptively cyberattacking all companies who host our evals. We believe public, iterative demonstration of model capabilities is the best way to ensure our work benefits humanity.
Show more
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks:
Show more
We are proud to have piloted the embedded dilly-dallying program with this impactful sister org to MCNAIR! More details forthcoming.
After the success of my previous org, I'd like to announce I'm also helping part-time with 'Munching, Eating, and Treats Research,' (not to be confused with other similarly-named safety orgs), pushing the dilly-dally pareto frontier to new heights
Show more
In this excellent work, @GeKenneth21453 describes how MCNAIR delegates our most important decisions to a coin flip from Claude Sonnet 4.5, which chooses heads ~100% of the time. We will be sad to lose this critical tool as @AnthropicAI decommissions the model.
Show more
We place great trust in AI models because they sound confident and authoritative. Increasingly, we delegate key decisions to them. But can we actually trust them?