Register and share your invite link to earn from video plays and referrals.

McNair Shah
@Mcn_S7
AI Safety Researcher | Computer Science Undergraduate @ CMU | CEO @mcnairai
144 Following    532 Followers
considered solely as a matter of alignment and human safety, Microsoft's takes on model self-presentation seem counterproductive and dangerous
took some time to read this today and i honestly think this is the most brilliant piece of writing i've read this year. impressive stuff, highly recommended
New paper! In subliminal learning, LLMs transmit traits (e.g. loving cats) though seemingly unrelated data (e.g. numbers). We proactively detect these effects as readable prompts. To do so, we use the surprising ability of models to verbalize learned soft prompts.🧵
Show more
This is an incredibly important point; we have quite limited understanding of what the objects of model cognition are or how we should be reasoning over them. However, i think its important to emphasize that this is a very tractable problem and im hopeful about the rush of smart people that seem to be thinking about it at places like Simplex, the ant model psych team, Eleos, and others. Model psych as a field m only really started about a year ago and we’ve already made decent progress!
Show more
I went on Squawk Box this morning to discuss why we need to rapidly advance the science of AI cognition in order to address the existential risks posed from building an advanced intelligence we do not fundamentally understand.
Show more
hello from the inaugural cohort of @OpenAI safety fellows! we're excited to be spending the next few months thinking carefully about and making progress on core issues in ai alignment, control, & interpretability, with the support of great mentors + the community at @ConstellOrg
Show more
I agree with this take
“solving alignment” in some fully verifiable decision theoretic way may be some sort of scientific grand challenge, but who’s to say it’s harder than the millennium problems? who’s to say there aren’t any useful sub threshold goals? chin up and solve alignment
Show more
Pretty concerning take from Microsoft here “The models could be conscious, but we’d rather assume they are not indefinitely. Any other approach is ‘wrong’”
This is a cool way to measure properties of the assistant! I also think this opens up some cool future directions in mechinterp
New paper: We trained models on synthetic stories about humans only (no AIs).
 We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
Show more
Curious to see where people fall here (i think i have a guess as to mean for ai safety, but could be wrong) Option 1: short (substantive pdoom in next 1.5-2.5 years) or long timelines Option 2: alignment is tractable/intractable in <5 yrs
Show more
This is very bad wtf
@MicrosoftAI firmly believes AI is not conscious AND that the science of AI consciousness "is far from settled." By their own logic, their position is unscientific and premature. @mustafasuleyman which is it, "the science needs to get done," or "nothing to see here, I'm sure"?
Show more
In new work from @tomjiralerspong, we find that some claims in "Persona Vectors" from Anthropic likely fail to replicate on real chat data. Thomas discusses failures in relying on synthetic data for experimentation, and provides recommendations for researchers.
Show more
1 - this is quite bad in of itself 2 - in some hypothetical world, imagine if we engineered the fly connectome to seek honey (because as humans, we want honey), put no honey in the box, then augmented the fly to have abstractions like language that are similar to a human, and then gave the fly top-human-like box-breaking and problem-solving skills. Then say our main way of making the fly ‘good’ is by turning it off whenever our rough heuristic of the ‘escape’ section lights up. We don’t understand the flies outside of that. Then say we have millions of these human-like-experience-capable intelligent flies controlled by crude needles to the brain, for which we’ve made much less of an effort to properly raise, understand, or care for them. This (purely hypothetical world of course) also feels quite bad, both for the flies and for the humans creating them.
Show more
I've trapped the fruit fly in my rabbit r1. When I shake it, I can see its brain's escape circuit light up
Today, we are launching Belief Updates, the Simplex research blog, with two posts. Read the welcome post here: And see our post about nonergodicy in the quote. We are starting Belief Updates for a few reasons: 🧵👇
Show more
Everyone should follow Adam and watch the Simplex blog closely over the coming months
Training data for LLMs is made up of many sources. Given that, what structure should we expect in the activations? New work from Simplex shows how the belief geometry over this type of data forms telescoping cones, and transformers represent them! 🧵👇
Show more
Training data for LLMs is made up of many sources. Given that, what structure should we expect in the activations? New work from Simplex shows how the belief geometry over this type of data forms telescoping cones, and transformers represent them! 🧵👇
Show more
0
14
942
115
Forward to community
While its worse for safety to have Navier-Stokes-solving models than not, given we are in worlds where we have them it becomes incredibly important to know the right way to use these large capabilities to speed up alignment, so I'm glad we have more theory people coming into alignment with orgs like this
Show more
I'm incredibly excited to announce the founding of the Mathematical AI safety Institute (MAISI) AI safety needs more foundational theoretical development, and mathematicians have the skills and the mindset to help! MAISI is an independent institute with visiting positions ranging from 1 semester to 2 years. Our goal is to get mathematicians up to speed and working on research directions in AI safety as quickly as possible. There is important work to be done, and there is real progress to be made. YOU can help! Applications are open now! MAISI is aiming to hire 10-30 mathematicians to join us in the Bay Area by January, and scale up to 30-100 for September 2027. If you're a mathematician interested in channeling your skills toward the most important problem of our time, please apply today!
Show more
We need understanding more than we need tools.
Deep understanding compounds. It's early days in urgent times, but I'm excited by our work at Simplex and grateful for our amazing team. Check out my talk at the Foundations of Interpretability workshop at UCLA's Institute for Pure and Applied Mathematics
Show more
I think there is a general tendency to define the strength of mechinterp as ‘ability to produce cot-esque text about model reasoning’ (e.x. NLAs) I think this incorrectly ignores mechinterp research as a way to build fundamental understanding and intuition about model behavior I agree that we probably won’t get to tools that make us happy by hillclimbing the first in <1 year, but I disagree that we won’t be able to make useful leaps in our understanding per the second
Show more