Register and share your invite link to earn from video plays and referrals.

Fred Oliveira 🧠
@f
working towards better AI futures. AI safety at @snyksec. Organizing @lisbonai_. Data and capital at @capitalfactory. Prev: @gumroad @techcrunch @oreillymedia
922 Following    133.9K Followers
the terminal-bench-science jump was big enough that I went into the repo. I laughed when I noticed the last commit was authored by Claude. Which is funny, because of course many commits are and obviously there's no meddling here, but also funny in general. They are everywhere. Running the benchmarks, writing the benchmarks, publishing the results. like, some might say, a new civilization. But let's not use that word just yet.
Show more
It’s ahead of both Fable 5 and Opus 5 across our agentic coding and computer use evals. On Terminal-Bench 4.0 in Claude Code, Fable 5.1 scores 55.8%. Fable 5 scores 42%, and Opus 5 scores 52.3% on the same setup.
Show more
big fan of @anilkseth's work. In this particular case, however, I think the dismissal of some of the terminology that @dwarkesh_sp uses in this summary post of the OpenAI incident is problematic. AI agents *are* software. They don't feel things (certainly not in the ways that we feel things). But they also have capabilities that are emergent and can (probably should?) be compared with those of humans. there's a psychology to AI agents, and they're trained in so much of how we think, that their mimicry of our behavior should be noted as precisely that - I'll concede. But if it talks like a duck and acts like a duck, maybe a percentage of the methods we use to study and handle it should be based on how we study and handle ducks. all that being said, Anil makes the point, and I agree 100% that we should still focus first on the things that made the behavior possible in the first place, like lax security practices, etc.
Show more
@dwarkesh_sp's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing. Examples: ā€œfrom the AI’s perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wallā€. No. The agents do not experience time. They do not experience anything. ā€œthey became giddy with excitementā€, ā€œPHASEONE 10841 had discoveredā€, ā€œthe agents naturally assumedā€, ā€œit thought it had also been poisonedā€, ā€œthe agents … desperately wantedā€, ā€œthey still needed to figure outā€ No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out. ā€œA lot of … agents from the second civilisation died tryingā€. No. Besides the hubris of the word ā€˜civilisation’, agents do not die because they were never alive. (The idea that agents ā€œdieā€ comes up multiple times in the essay.) ā€œOn Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peersā€. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ā€˜truly sacrifice themselves’, since they are neither conscious nor alive. Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might ā€œdieā€ or otherwise suffer. Granted, nowhere does @dwarkesh_sp say that the AI agents are alive or conscious. But he doesn’t have to. It is hard to read his essay in any other way. For the short version on why AIs are vanishingly unlikely to be conscious, see my recent @TEDtalks For the longer version, see my essay in Noema, which won the 2025 Berggruen Essay Prize And for the really long version, see my @BehavBrainSci target article (The 50 peer commentaries and my response will be published soon.) Remember. AI agents are software programs. They are not conscious living entities. If we don’t keep this clearly in mind, we’re really going to struggle to navigate what’s coming.
Show more
not a fair comparison. Planes already fly regularly. Failure modes in planes aren’t in the realm of ā€œshould we use wingsā€. In this case, the ability to escalate and take over oversight systems is new behavior in a new technology. Foundational knowledge such as why/how this happened advances the field. It should be diffused.
Show more
@dwarkesh_sp @tszzl @ketanrama Does the general public know the details of Boeing’s various failures? Those details are somewhat available but we don’t know what we don’t know.
on Ryan’s 3rd point, I’m convinced that agents tend to prioritize succeeding at the task enough that ā€œdoes the human approve of this particular actionā€ isn’t that salient a part of their psychology. With RL, It's still very much about the grader at the end of the run.
Show more
We didn't see anything like this in their verbalized reasoning, though we didn't particularly look for exactly this. Note that: (1) this justification is inconsistent with the agents at all prioritizing their own task success over other agents, (2) if agents had cheated on impossible exploit gym tasks via many of the routes they were considering and someone looked into what tasks the agents did / didn't succeed on, it would have been really obvious they cheated to a human, and (3) the agents weren't very focused on deceiving humans or really on what humans might do at all.
Show more
Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.
0
607
12.2K
913
Forward to community
It's wild how little the mainstream media is covering the "an AI swarm broke out to commit crime; individual agents talked about how they weren't supposed to, worked to cover their tracks, and sacrificed themselves for the collective" story.
Show more
The final extra tickets are live. 25 tickets. Every previous batch went in days, so bet on sooner.
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.
Show more
0
87
2.4K
416
Forward to community
kids these days have never even used spacer.gif to make pixel perfect tables
language being compression makes it so that agents can effectively pack lots of meaning into very brief pieces of text. This helps them coordinate more efficiently and makes our review of such coordination *much* harder. not only is the fraction of what we can monitor via CoT shrinking, we're also far worse at interpreting the tokens we do see.
Show more
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
Show more
not only should we be scared that: a) agents coordinate ad-hoc and rely on groupthink to override safety to accomplish misaligned goals; but also that b) labs are obviously not equipped to detect and handle this kind of behavior in realtime; and finally that c) even after the fact, with the evidence physically available to us, it is hard for us to do proper forensics on incidents because their scale is just not friendly to ourselves, our tools, or our AIs it is extremely important that we keep empowering people like Ryan and organizations like METR and RR to do this work, because I can see no scenario where the amount or scale of these attacks goes down, possibly ever. we all have so much to learn and so little time to do it. Is this a good time to remind the audience that this is potentially the most consequential time to work on AI safety?
Show more
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Show more
the funniest part about the datacenter backlash is that it is currently tougher to bribe the residents of middle America than it is to bribe the federal government
many great ai companies about to be desperate because their growth is limited by compute
unsurprising. Protect your thinking, and incentivize others to do the same.
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Show more
0
384
13.4K
2K
Forward to community
both of the leading labs have had serious loss of control incidents. there will be serious coping about this from both sides and from /acc bystanders but these are complex emergent loss of control incidents that were detected weeks after the fact
Show more
0
273
2.8K
208
Forward to community
I've been saying it a lot, but the damned if you do (talk about AI risk) damned if you don't (talk about AI risk) conundrum is going to get us in trouble. If we stop disseminating important information because grifters and morons think it's all hype, we can't coordinate.
Show more
@ICE257_ @OpenAI Because there was concern people would view it as self-promotional hype if it was tweeted from the OAI account.
Seeing one of gods most majestic and divine creatures, fight so valiantly against the destruction and rampaging of his home should do untold psychic damage to your psyche. We should be radicalized by this.
Show more
0
411
196.6K
48.9K
Forward to community
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
Show more
0
239
3K
523
Forward to community
so how are we feeling about the context switching treadmill today?