Register and share your invite link to earn from video plays and referrals.

Mark Saroufim
@marksaroufim
Fighting Amdahl @coreautoai
986 Following    17.2K Followers
The megakernel literature, summarized
It took me a long time to build an intuition for why CoT works. My thinking was always.. if the model can predict it downstream of 10k thinking tokens, it should have been able to predict it from the outset too. My intuition now is: - During inference, the correct paths are indeed somewhere in the hidden states, represented purely as probabilities - However, in the process of sampling, we're forced to materialize just one path. This is destructive -- a 30% chance of ending up at the answer can become 0 if we sample the wrong token. - The constant backtracking reasoning models do protect against this. Every "wait" or "but" is another chance for a shot on target. - By the time models exhaust their reasoning budget, they've already seen a bunch of possible answers - And since these models are also generally better at verifying answers than generating them, the chances of choosing the correct path, conditioned on this prefix, are much higher than it was at the start.
Show more
One Layer Deeper is officially live! The motivating idea is that some models just don’t want to learn. Most optimizer benchmarks ask how quickly you can train a given model, but baked into that question is the assumption that it can be trained (well) at all. We think this matters especially for adaptive computation. Ideally, a model should be able to spend (far) more compute on harder problems and use more computation at test time than it spent on any given training instance. There are many possible ways to do this, and we don’t want to assume the answer is one specific recurrent or looped architecture. Making this all work is really an architecture-optimizer co-design problem. Rather than fixing the optimizer, loss, and training setup, One Layer Deeper asks people to design all of these together. For the competition task, we use repeated modular squaring, x^(2^T) mod N. The task is inherently serial because each squaring needs the residue produced by the previous one. When N is a semiprime, the only known way to skip ahead requires knowing its factorization. Increasing T therefore adds another genuinely dependent step without making the input or output longer. We are very excited to work on this with @marksaroufim @_arohan_ and everyone on the @CoreAutoAI team. Submissions: Blog Post:
Show more
watching my loss not even move an inch on my 999th attempt at one layer deeper competition by @tilderesearch x @CoreAutoAI
AMD MI355X vLLM HAS BEATEN B200 vLLM ON KIMI K2.5 (THE SAME MODEL ARCH AS XAI CURSOR COMPOSER 2.5)🚀🚨 This uses upstream AMD kernels from the @GPU_MODE community. We explain below 👇 1/6🧵
Show more
It was an honor to give this talk especially at the OG YC building. I grew up on PG essays and it was exciting to meet so many young entrepreneurs working on hard systems problems.
At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence per watt, heterogeneous inference, and more. Thank you to our presenters: 0:00 – @FrancoisChauba1: The case for chip and kernel specialization 7:16 – @stuart_sul: Parallel Kittens - Systematic and Practical Simplification of Multi-GPU Al Kernels ( 21:29 – @JonSaadFalcon: Intelligence per Watt - Measuring the Intelligence Efficiency of Local and Cloud AI ( 31:05 – @MarkSaroufim: When Al Starts Writing Systems Code 47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware 1:04:33 – @shacklettbp: Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU (
Show more
GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON. The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵
Show more
The Ryzen Halo by @AMD is a work of art, it's quiet and cool, has a subtle but slick Need for Speed Underground aesthetic, blazing fast, most AI software worked OOB. Between this & the @FrameworkPuter desktop, AMD is shipping the most beautiful electronic devices on the market.
Show more
I stand on the shoulders of giants who stood on the shoulders of giants. So much of the progress we’ve made has been because of the openness of our great predecessors. Proud to support
Show more
A form of looped compute, but not serial compute — and therefore insufficient for reasoning. In our recent work, The Seriality Gap in Video Diffusion Models ( we show denoising steps do not provide scalable serial computation. Feel free to check it out!
Show more