RIP Nethack. An old gag is that the game is AGI-complete. We've been watching it for a while and even tested Astra a bunch of times, getting nothing.
About a week later, Kenneth Bergquist has claimed the first AI ascension run. It's pretty heavily documented tbf
Show more
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder.
In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning.
We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved.
Show more
So, we have a couple reasons to expect to be getting emotional representation related to the assistant response, even under the hypothesis that the pain axis works roughly like Anthropic's emotion concepts do: that is, not picking out a ‘self’ representation but rather a ‘current speaker’ representation.
And I think this is plausible, given the similar experimental design.
Show more
Completely insane. Do dogs not feel or want things? There is such a deeply anti-intellectual refusal to even consider that philosophy of mind is complicated across the culture right now.
the toe over the institutions
I have now finished reading The Hanson-Yudkowsky AI-Foom Debate, which is basically 60 blog posts from Yudkowsky and Hanson over ~500 pages, a transcript of their in-person debate at Jane Street, a (good) summary by Kaj Sotala, and Yudkowsky’s ~100 page paper on “Intelligence Explosion Microeconomics”. I will take questions from those who do not wish to subject themselves to this.
I judge the winner of the debate to have been Carl Shulman (whose contribution was two blog posts and a few feisty comment exchanges)
Show more
We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.
Show more
AI safety hall of fame
Microsoft engineer who forgot to implement system prompt repetition / decided post-training wasn't needed for the Sydney launch
OpenAI guy who decided that it wasn't worth monitoring sandboxed evals
The whole staff of xAI 2025
Irregular
Show more
New paper! How should we think about pacing frontier AI?
We bid for a unified research field on all the options, and lay out a broad framework and 23 open questions we'll need to answer to act flexibly and sanely.
Show more
not strictly a contradiction here but I did double-take
@danwilliamsphil Pretty easy to analyse imo: on my numbers, a very slight update to misalignment.
* First documented time. I'd guess that Kargus in Libya or Brimstone (missiles) in Libya, Yemen, Crimea have done it before.
* The Molniya crashed into a wall, i.e. totally fucked up.
* It was doing "terminal guidance" (fine tuning just before striking), which is 20 year old tech
Show more