@willdepue That's exactly what I have told everybody since the 1970s
In 1971, the deep learning pioneer Ivakhnenko modeled the British economy through a deep 8-layer network. He didn't use the new term "Artificial Intelligence;" he still used "Cybernetics." See the most accurate history of deep learning & cybernetics & modern AI:
Show more
Today everyone is talking about Recursive Self-Improvement (RSI). In 1987, when compute was 100,000,000 x more expensive, I published the 1st concrete RSI algorithms. Now compute is cheap, and RSI is driving the future of both software and physical AI. See: RSI since 1987 (Technical Note IDSIA-9-26)
Also covered: RSI with self-modifying policies since 1994, gradient descent-based RSI in neural networks since 1992, asymptotically optimal RSI for curriculum learning since 2002, mathematically optimal RSI through the self-referential Gödel Machine since 2003, RSI combined with artificial curiosity and intrinsic motivation since 1990/1997, recent work on RSI since 2020.
Software-based RSI has become practical. Full RSI, however, will require not just self-improving software but self-improving hardware in the physical world.
As of 2026, companies talking about RSI include Anthropic, OpenAI, Sakana AI, SpaceX, Ricursive, Recursive Superintelligence, Inherent …
Show more
Current LLMs aren't truly creative because they haven't implemented the "Formal Theory of Fun and Creativity" (2008) yet. See "Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes" Tweet:
Show more
Certain companies train their AIs to plagiarise work by unnamed scholars, extending an AI history permeated by plagiarism and misattributions. It started a long time ago. In the 1900s, the old method of least squares (Gauss & Legendre, 1795-1805) was renamed "neural network," without citation of the original. In the 1950s, cybernetics was renamed “AI" by people who did not want to credit earlier AI pioneers such as Wiener. In the 2000s, AI techniques of the 1960s (Ivakhnenko & Lapa 1965, Amari 1967) were renamed “deep learning” without correct attribution. I could point out many additional cases. Some of the plagiarists even got awards. In the best interests of the field, and scientific honesty in general, it is time to stop this.
Alas, in the end, the facts must always win. As long as the facts have not yet won, it's not yet the end. As Elvis put it, "Truth is like the sun. You can shut it out for a time, but it ain't goin' away.”
More tweets and reports on this:
Show more
This "recurrent depth" is essentially what's in Sec. 5.3 of the 2015 paper: On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models This paper went beyond the inefficient millisecond by millisecond planning of my 1990 neural world models, addressing planning and reasoning in abstract concept spaces. The 2015 control network C is a prompt engineer that learns to create a chain of thought: to speed up decision making, C learns to query its separate neural world model for abstract reasoning. The prompts and the answers are internal self-generated sequences of vectors that don't have to represent natural language.
Show more
The first modern backprop-trained CNN for vision was published in 1988 in Japan by Wei Zhang (a Chinese researcher), J. Tanida, K. Itoh, Y. Ichioka. Wei brought it to Silicon Valley, earning 1998 FDA PMA approval for radiology’s 1st AI, reading 10M+ mammograms/yr in the US.
Show more
I support open-source models distilling what commercial companies distilled for free from the entire internet. I published distillation for free in 1991 in Europe - this was copied in the US and in China (
Show more
I am proud of the work my team did in Munich in 1991, when compute was millions of times more expensive. We published the roots of today's trillion-dollar AI boom:
★ 3/1991: the first kind of Transformer (see the T in ChatGPT) - now called the unnormalized linear Transformer: the predecessor of the normalized quadratic Transformer
★ 4/1991: Pre-Training (the P in ChatGPT) & Neural Net Distillation (see DeepSeek and many other LLMs)
★ 6/1991: Deep Residual Learning, basis of LSTM & Highway Net / ResNet (most-cited AIs of their centuries)
★ 8/1991: conference paper on GANs for World Models trained by Artificial Curiosity
★ Around the same time, Munich also was the origin of the first self-driving cars in traffic (Ernst Dickmanns et al.), going up to 175 km/h. The city was truly the epicenter of AI.
Read the timeline with links to the original references, featuring a preface by
@hardmaru:
Show more