가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Corey J. Gallon
@CoreyGallon
Sharing insights from the frontiers of AI Engineering. 🇦🇺 Technologist. Investor. Coffee Nerd.
가입 December 2008
319 팔로잉 중    544 팬
Every axis of scaling we have has only ever been pointed at public data. Wikipedia, Reddit, arXiv, GitHub. None of it is pointed at your emails, your meeting transcripts, or the work your company actually does. @jxmnop, cofounder at Engram, spends "Scaling Compute on Context" on that gap, and the talk is on @aiDotEngineer's YouTube. It's a tour of the methods people are trying for getting a pre-trained model to know your private data, and where each one runs out. - Breadth without depth. Terence Tao's point about AI: it knows every public mathematical topic and can connect them in ways no person could, but it lacks the intuition a grad student builds over five years in one area. - Two of the three scaling axes are closed to you. You scale by adding data, adding compute, or growing the model. With a fixed private corpus you can't make more data and you won't train from scratch, so compute is the one you have left. - Next-token training on your own corpus collapses the model. Take 10K financial reports, drive the loss to 0.0001, and generation falls apart. It also can't answer a question unless the answer sits in the data already. - Compaction buys context, not gradients. Compressing the corpus into a small set of KVs, the way Claude Code and Codex compact, only covers what fits in context and skips what taking gradients gives you. - On-policy distillation trains the model to act as if the data were in context. Raw documents don't distill well, so the self-study approach in the cartridges paper generates question and answer pairs conditioned on the corpus first. - Synthetic continued pre-training is promising and awkward. It overwrites part of the original pre-training, and it wants a base model, so you're post-training all over again afterwards. - All of these hit a wall. You define a data set, you train, you fit it, and then more compute stops buying more depth. - Self-improvement is the missing piece. AlphaGo got better by making its own training problems harder as it improved. The curve worth chasing is one where the model keeps generating harder data for itself. - The name isn't settled. Sleep-time compute, continual learning, neural memory, note taking, machine studying, amortized inference. One idea under a pile of names, because the paradigm is early. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
더 보기