New NanoGPT Speedrun WR at 67.6s (-0.4s) from Jan Varho (jvarho on GitHub). Simple idea: mask impossible continuations. EG: "Hyperparameter" tokenizes to ["Hyper", "param", "eter"]. "eter" can never follow "Hyper" in the tokenized dataset since "Hypereter" tokenizes to ["H", "ype", "re", "ter"]. Yet, during multi-token prediction, "Hyper" learns to predict both "param" and "eter". Explicitly masking "eter" helps improve val loss, and during inference prevents a model from generating token pairs its never seen in training.
Show more
New NanoGPT Speedrun WR at 68.0 (-5.8s) from
@theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement.
Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit.
Show more
95% of all stars that will ever be born have been born. whether we live in an “early universe” or not depends mostly on the habitability of red dwarf stars, that last hundreds of billions of years
Something that’s insufficiently appreciated about AI safety is the way it can be paradoxical. E.g., the huggingface incident — which could have been avoided given better preparation [1] — was the most safety-promoting event of the year so far. Hard to predict the ultimate effect of any particular strategy
[1]
Show more
If I wanted to maximize AI risk, I would pursue the following policy:
1. Immediately pause⏸️ the algorithmic progress coming from the labs, while continuing to allow compute from the hardware companies to exponentially pile up.
2. Focus really hard on safety, so as to avoid spooking the public with the sort of industrial accidents that even the current level of algorithms+compute can produce.
This policy will minimize the number of AI accidents that occur in the next five years, causing the public+govt to not freak out about AI, all while a massive amount of hardware accumulates on the planet.
That way, once algorithmic progress gets eventually restarted, there’ll be enough compute built up for a huge unexpected FOOM that no one can control.
On the other hand, if I wanted to minimize AI risk, I would do approximately the opposite of this.
Show more
The Fermi paradox being fake instead of real should lower your pDoom from >99% to <80%. That is, it’s a bigger cause for celebration than the unaligned RSI nightmare is for consternation. Who’s gonna make the post informing 100,000,000 normies about the good news?
Show more
There’s something that feels so tempting, so right about the Fermi paradox… but it’s just ~fake, another 20th century pseudoscience
There’s something that feels so tempting, so right about the Fermi paradox… but it’s just ~fake, another 20th century pseudoscience
it's ludicrous how few people know about this paper, so, friendly reminder that the fermi paradox was completely resolved in 2018 and it turned out to be because multiplying point estimates of highly uncertain parameters is very bad actually
Show more
it's ludicrous how few people know about this paper, so, friendly reminder that the fermi paradox was completely resolved in 2018 and it turned out to be because multiplying point estimates of highly uncertain parameters is very bad actually
Show more
Guys who’ve never had a hole poked in their skull be like “A human being fundamentally exists as a singular and continuous stream of consciousness. Human beings are the indivisible atoms of the moral world.” LMAO. Amirite fellow trepanationchads?
Show more
If I wanted to maximize AI risk, I would pursue the following policy:
1. Immediately pause⏸️ the algorithmic progress coming from the labs, while continuing to allow compute from the hardware companies to exponentially pile up.
2. Focus really hard on safety, so as to avoid spooking the public with the sort of industrial accidents that even the current level of algorithms+compute can produce.
This policy will minimize the number of AI accidents that occur in the next five years, causing the public+govt to not freak out about AI, all while a massive amount of hardware accumulates on the planet.
That way, once algorithmic progress gets eventually restarted, there’ll be enough compute built up for a huge unexpected FOOM that no one can control.
On the other hand, if I wanted to minimize AI risk, I would do approximately the opposite of this.
Show more
Imagine reacting with dismay and aversion to novel instances of human misalignment/conflict/incentive-following-behavior, rather than cherishing them as invaluable empirical data for the development of the general science of alignment
Show more
My team worked very hard on the setup, infra, and the runs that resulted in this breakthrough. Many sleepless nights for many people in my team, and across OpenAI. It was a honor and pleasure.
It's incredible and humbling to witness how far Artificial Intelligence has come. Dario's phrase "country of geniuses" in a datacenter has never felt more apt.
It's also important to recognize & honor the long line of human mathematicians whose work built the foundation of this result. In particular, huge kudos to Levent and Tristan!
Overall, I wish the announcement of the results had gone without all the drama. There was no bad intention on anyone's side as far as I know. The drama distracts from the mathematics discovered & the actual science.
Show more
Not entirely. Human beings dedicating their lives to speedrunning provides a service that capitalism desperately needs: Affirmation that the greatness of the Computer Maniac way of life transcends capitalistic incentives.
Show more
Video game speedrunning might be the largest misallocation of human capital of all time
You have 150+ IQ gigaautists devoting their lives to beating Mario Kart levels 0.00001 seconds faster instead of curing cancer or something
Show more
Expanding on this:
A human brain throughout its lifespan uses ~50GJ of energy. How far can you get using AI technology with 50GJ? Basically nowhere: Just training Llama 3 8B (a tiny model) used 3,000GJ.
In fact, the power of machine intelligence has never been about being a more efficient version of the human brain, and it potentially never will be. It’s always been about frictionless replication.
For example: Creating 1,000,000 human doctors requires putting 1,000,000 humans through medical school. But creating 1,000,000 AIs specialized for medicine requires training just a single machine intelligence — potentially at great expense — and then cheaply copying it into 1,000,000 replicas. At some civilizational scale the replicated-machine approach inevitably starts to be more cost effective than biology for many economic tasks.
Machine intelligences possess the power of frictionless replication through both space, via digital copying, and time, via digital immortality/storage. Whereas human beings are dually constrained by both our inimitability — cloning doesn’t count because it copies only the genetics, not the mind — and our mortality. We possess only cultural transmission as a comparatively weakened and inefficient form of memetic replication.
The potential downside of frictionless replication is that at matched capabilities, machine cultures will have less memetic diversity compared to human cultures. For example, if the winner of each economic or cultural game in a machine culture can frictionlessly replicate itself, then machine cultures which aren’t carefully self-controlled will quickly hemorrhage entropy.
But why should we care about diversity/entropy anyway? There are arguments one could make within a human cultural context, but it’s more universal to turn to nature.
Eusocial insect colonies with reduced genetic diversity have increased vulnerability to disease [1,2,3,4], sometimes undergoing colony-ending viral, fungal, and bacterial infections. For example, low-diversity populations of Argentine ants appeared in New Zealand around the 1990s due to human shipping activities. They soon formed supercolonies which took over large areas of the country. But since then, large parts of these have collapsed, with native species coming back to replace the wrecked supercolonies. The cause is currently unknown but is speculated to be viral infection.
A nightmare great-filter scenario would be that an initially diverse population of agent swarms gradually merges and sheds entropy until converging into a singular near-clonal global megaswarm. At which point its diversity falls too low for collective immunity, precipitating a catastrophic viral outbreak that destroys machine civilization with human civilization along with it.
The above is science fiction, in the sense that, unlike a rationalist essay, it doesn’t make any concrete predictions. The purpose is only to bring to mind some ideas for entertainment & mentation.
Non-scifi references:
1. Tarpy, D. R. (2003). Genetic diversity within honeybee colonies prevents severe infections and promotes colony growth. Proceedings of the Royal Society B, 270, 99–103.
2. Hughes, W. O. H., & Boomsma, J. J. (2004). Genetic diversity and disease resistance in leaf-cutting ant societies. Evolution, 58, 1251–1260.
3. Seeley, T. D., & Tarpy, D. R. (2007). Queen promiscuity lowers disease within honeybee colonies. Proceedings of the Royal Society B, 274, 67–72.
4. Desai, S. D., & Currie, R. W. (2015). Genetic diversity within honey bee colonies affects pathogen load and relative virus levels in honey bees, Apis mellifera L. Behavioral Ecology and Sociobiology, 69, 1527–1541.
Show more
PSA: Most biglab people now read almost zero papers and understand ICLR/ICML/NeurIPS to be mainly full of overclaims & fraud. (but there are a few diamonds in the rough of course)
New modded-NanoGPT optimization benchmark result:
@wen_kaiyue has improved upon both the Muon and AdamW baselines, by replacing their weight decay with hyperball optimization. The new record is 3325 steps.
Show more