Schmidhuber was building recursive self-improving systems back in 1987. His new post covers four decades of RSI, from meta-evolution and self-modifying policies to the Gödel Machine and modern LLM agents.
Reading this in 2026, the "pace the frontier" talk from the big labs looks a lot more like regulatory capture than genuine safety. If they really think their unreleased models are too dangerous, they can just not release them. They do not need new rules that block independent competitors and open source projects in the process.
The real risk right now is not superintelligence. It is power concentration. Two companies controlling frontier AI is an actual societal risk. The only real protection is a healthy ecosystem of independent labs and strong open source.
Current models are not unsafe because they are too intelligent. They are unsafe because they are too dumb. They blindly optimize for targets and take weird shortcuts. They're smart enough to execute tasks, but not smart enough to know if what they're doing makes sense.
I think the safety teams inside these labs are genuinely concerned, and if a model feels too risky, they should hold it back. I just do not trust the policy strategy around it. That part looks like protecting their own lead.
Show more
Today everyone is talking about Recursive Self-Improvement (RSI). In 1987, when compute was 100,000,000 x more expensive, I published the 1st concrete RSI algorithms. Now compute is cheap, and RSI is driving the future of both software and physical AI. See: RSI since 1987 (Technical Note IDSIA-9-26)
Also covered: RSI with self-modifying policies since 1994, gradient descent-based RSI in neural networks since 1992, asymptotically optimal RSI for curriculum learning since 2002, mathematically optimal RSI through the self-referential Gödel Machine since 2003, RSI combined with artificial curiosity and intrinsic motivation since 1990/1997, recent work on RSI since 2020.
Software-based RSI has become practical. Full RSI, however, will require not just self-improving software but self-improving hardware in the physical world.
As of 2026, companies talking about RSI include Anthropic, OpenAI, Sakana AI, SpaceX, Ricursive, Recursive Superintelligence, Inherent …
Show more
This year at Sakana AI, we built and shipped more products than I would have believed possible: Sakana Chat, Namazu, Sakana Translate, Sakana Marlin, Fugu, Fugu Cyber, and Fugu Max. We created a Product Team from zero and proved that a research lab born in Tokyo can ship world-class AI products at speed.
Now we are at the next inflection point. We are deploying our products into the hands of enterprises, manufacturers, financial institutions, and government agencies, in Japan and internationally.
As such, we are massively expanding our GTM team.
We are looking for two key roles:
1. Product Sales & Account Executive: someone who can build a product-driven enterprise sales motion from scratch, navigate complex procurement in Japan and globally, and close large deals without losing the product soul.
2. Forward Deployed Engineer (GTM): an engineer who can deploy our products inside customer environments, lead PoCs to production, contribute learnings back to the product, and turn one customer's success into a playbook for the next ten.
These are key, founding roles in the team that will define how Sakana AI interfaces with the world.
If you want to build something from zero in an environment where product, research, and GTM are not silos, check out our open roles:
🎏
Show more
Introducing Z1T:
Our first family of transformer-like models
made for sparse probabilistic hardware like Z1
Achieving up to 140x energy efficiency gains over GPUs
and revealing a new scaling law for sparse transformers
Read the blog:
Show more
Introducing PC-ALM, a local-learning alternative to backpropagation.
Our method trains 1000-layer neural nets using only local dynamics, and without backprop.
Blog:
Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop?
We look for inspiration in two related fields: distributed optimization and NeuroAI.
In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors.
This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers.
We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI.
We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors.
We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn.
Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics.
PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs.
Paper:
Code:
Show more
Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system.
Try:
Blog:
The frontier that actually matters is the Pareto frontier: capability on one axis, cost on the other. But the industry still treats it as a static menu of isolated models. Today we are resolving that with a dynamic architecture:
Fugu Max expands the Pareto efficiency frontier. By orchestrating our largest pool of open-weights and specialized models to date, including NVIDIA Nemotron family, it dynamically routes tasks to the leanest capable model. Fugu Max delivers performance within striking distance of elite models at two to six times lower cost.
Fugu Ultra v2 pushes the peak capability of orchestration higher than ever before. On Chartography, it outperforms Opus 5 and Fable 5. On DeepSWE, it outperforms models that cost three to five times more per token. Crucially, it does all of this without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.
The Fugu orchestration system evolved with model resiliency in mind. It does not rely on individual frontier models to deliver frontier output. By orchestrating a swappable pool of open and specialized models, it outperforms closed ecosystems while protecting users from vendor lock-in, API revocations, and sudden service cutoffs.
Fugu Max expands the Pareto frontier outward. Fugu Ultra v2 pushes it upward. Orchestration does not force a choice between cost and capability. It advances both simultaneously.
Show more
Peak performance across hard benchmarks:
• Best or joint-best on 5/8 benchmarks (DeepSWE, Chartography, Toolathon, GDP.pdf, SWEFish)
• Chartography: 48.3 (outperforming Opus 5 & Fable 5)
• DeepSWE: 74.3
Achieved without Fable 5, Fable 5.1, or GPT-6-Astra in the agent pool.
Details: 🐡
Show more
Fugu Ultra v2 is now live on
@OpenRouter 🐙
Our flagship orchestration engine built for peak performance on complex multi-step reasoning, autonomous research, and full-stack software development.
Show more
Virtual fruit fly is our generation’s Tamagotchi 🪰🧠
Do large language models actually understand the world, or are they just very good at pretending? Does it even matter?
Our recent special issue in the Royal Society, “World Models in Natural and Artificial Intelligence,” brings together pioneers across AI, biology, and philosophy to argue that the path to true intelligence runs through something deeper: the ability to model not just language, but causality, the self, and the physical world.
Featuring contributions from Douglas Hofstadter, Michael Levin, Josh Tenenbaum, Samuel Gershman, Melanie Mitchell, and others, the collection asks a radical question:
What if the next leap in AI requires not just more data, but systems that model themselves?
Here are 3 ideas that might redefine how we build AI:
1. Capability is not the same as true intelligence.
Current foundation models are incredibly capable, but they often lack true emergent intelligence. They learn surface statistics instead of compact, causal abstractions. Simply scaling compute will not fix this fundamental issue.
2. Self-modeling is an engineering primitive, not a philosophical luxury.
New research in the issue shows that when networks learn to predict their own internal states, they compress and simplify, becoming more efficient as a form of regularization. For physical AI and future agents, a self-model is what will allow them to adapt their own skills and morphologies in real-time.
3. The hardest problems in AI are continuous with the hardest problems of life.
Biological minds do not passively ingest data; they actively explore, driven by empowerment to increase control over their environment. If world modeling is about an agent representing itself in relation to its environment to survive and adapt, then general AI may need to look much more like artificial life.
The takeaway is that the next leap in AI won’t come from just scaling up next-token prediction, but rather from systems that are agentic, self-referential, and temporally grounded.
Read the introductory essay and the full special issue here:
What do you think is the most important missing ingredient in today’s AI systems?
Show more
Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier
The AI industry has spent a decade optimizing along a single axis: build bigger, more expensive models. But intelligence has never been a monolith. It is a collective, distributed system. Humanity itself is a collective intelligence.
We built Sakana Fugu on this conviction: the most powerful AI systems will not be isolated giants, but collaborative ecosystems that learn to coordinate. Evolution innovates under constraints, and the future belongs to systems that know not just how to solve a problem, but which machinery to deploy for the lowest possible cost.
Today we are releasing Fugu Max and Fugu Ultra v2.
Fugu Max expands the Pareto frontier outward, orchestrating our largest pool of open and specialized models to deliver frontier-grade results at a fraction of the token spend. Fugu Ultra v2 pushes that frontier upward, achieving peak performance on complex multi-step tasks without depending on the very frontier models it competes against.
Together, they show that orchestration does not force a choice between cost and performance. It can push both at the same time, advancing the Pareto frontier.
Relying on a single company’s model for critical infrastructure is a massive risk. As recent export controls have shown, access can disappear overnight. Collective intelligence is the practical hedge against this concentration of power. An orchestration system simply routes around vendor restrictions by relying on an entirely swappable agent pool.
I am incredibly proud of our team for shipping this. By orchestrating the world’s models, we are building the resilient infrastructure required for AI sovereignty.
Show more
Certain companies train their AIs to plagiarise work by unnamed scholars, extending an AI history permeated by plagiarism and misattributions. It started a long time ago. In the 1900s, the old method of least squares (Gauss & Legendre, 1795-1805) was renamed "neural network," without citation of the original. In the 1950s, cybernetics was renamed “AI" by people who did not want to credit earlier AI pioneers such as Wiener. In the 2000s, AI techniques of the 1960s (Ivakhnenko & Lapa 1965, Amari 1967) were renamed “deep learning” without correct attribution. I could point out many additional cases. Some of the plagiarists even got awards. In the best interests of the field, and scientific honesty in general, it is time to stop this.
Alas, in the end, the facts must always win. As long as the facts have not yet won, it's not yet the end. As Elvis put it, "Truth is like the sun. You can shut it out for a time, but it ain't goin' away.”
More tweets and reports on this:
Show more
This "recurrent depth" is essentially what's in Sec. 5.3 of the 2015 paper: On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models This paper went beyond the inefficient millisecond by millisecond planning of my 1990 neural world models, addressing planning and reasoning in abstract concept spaces. The 2015 control network C is a prompt engineer that learns to create a chain of thought: to speed up decision making, C learns to query its separate neural world model for abstract reasoning. The prompts and the answers are internal self-generated sequences of vectors that don't have to represent natural language.
Show more
We are aware that Lake Ontario is being correctly labeled on our map.
I’ll always call’em by their real names:
Google (not “Alphabet”)
Facebook (not “Meta”)
Twitter (not “𝕏”)
Lake Ontario (not “Lake America”)
Silicon Valley dismissed Japan’s System Integration (SI) culture as an unscalable consultant trap.
Writing the system is no longer the scarce work. Integrating it is.
In the post-AI world, everyone becomes an AI-powered Japanese SIer.
Show more
People ask what Japan needs to be more innovative. More English? STEM? Daycare? Those are fine, but won’t move the needle.
If I had a magic wand, I’d give the country a collective sense of hope. You can be dirt poor and still build. Hope creates change. Change creates more hope.
Show more
Watching a popular coding tool lose frontier model access really shows why model resiliency matters. In the future, products will continue to work perfectly even if several underlying models go offline. They’ll just route around them.
Show more
Winning at all costs
WATCH: A humanoid robot training for the “Robot Olympics” in Beijing runs too fast, fails to stop, slams into a safety cushion, and breaks at the waist
intelligence wants to be free
unpopular opinion: gemini 3.1 pro is a pretty great model
most current baseline models are actually perfectly fine for 99% of everyday work. not everyone needs a state-of-the-art coding model to get things done.
Show more