Register and share your invite link to earn from video plays and referrals.

Jack Morris
@jxmnop
research @engramlab // language models, information theory, science of AI
1.1K Following    53.7K Followers
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing. - pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper - use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this - use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs" "De-identification" is weak -- you can identify someone with a small number of bits, and long traces have more than enough. And it doesn't affect IP leakage concerns.
Show more
0
44
1.4K
154
Forward to community
this is insane, if they named their company ☺️ back in 2016 they would only be getting acquired for 978 million
JUST IN: NVIDIA's $12,930,300,000 acquisition of Hugging Face contains an easter egg. The number 129,303 is the decimal conversion of Unicode point U+1F917. The 🤗 emoji.
who told my local restaurant about claude
0
129
12.3K
268
Forward to community
Question for LLM enthusiasts these days, one typically kicks off post training by distilling some reasoning trajectories from some other model. it's kind of like a sourdough starter so how was the *first* model post-trained? did humans write the original set of formatting traces from scratch? (this is why Inkling used a small number of Kimi traces btw)
Show more
Just realized that artificial intelligence is the future
very cool idea. but don’t none of you dare trying to take this cute innocent little model and use it for your large-scale GRPO. its not just in bad taste, it violates international child labor laws
Show more
What happens when an LLM never sees material beyond fifth grade? We trained a 5B LittleLearner model from scratch on LittleCurriculum, a corpus restricted to K–5 material, to make this question testable.🧵 💬 Chat with LittleLearner yourself:
Show more
i'm quite excited about this! • it's cool in general that models can generate their own training data and learn from it. wasn't fully clear to me even a year ago that this would work reliably. • the thing that we're trying to build seems really important and no one has built it before • some traces from our model genuinely surprise me (e.g. the attached example, where it perfectly simulates the output of a complicated bash query, despite not being trained to do this) my main feeling overall is that there is so much to do. we see signs things are starting to work: models use their memories to produce better answers, knowing things saves lots of tokens, and all of this emerges with more compute but solving this memory calibration problem, on top of learning how to generate data in ways that scale nicely with compute, is going to take some time this is just a first step 🫡
Show more
Almost got run over on my bike yesterday directly outside the xAI office. i look up and it's a special black tesla with license plate MINLOSS. just incredibly on brand. no notes
i've started having claude turn my codebases into visual diagrams so i can discuss the codebases with claude more easily - the moving dots are data snippets that i can inspect
0
185
7.2K
296
Forward to community
If you think about it, it's pretty embarrassing that frontier labs still pretrain models. Don't they know there's a theory for this? Given how big the model is, and how long you're going to train, you can predict the loss exactly. No need to actually train. They spend billions on this!
Show more
But seriously. SF housing do costs seem crazy. Why are people willing to pay that much? Like, why would so many people be willing to spend so much to have great weather near two great universities and proximity to the most dynamic companies in the hottest industry in the world?
Show more
0
89
1.1K
30
Forward to community
I used to walk around thinking about code. I could "see" the code in my head. It was pretty cool. If there was a bug I could often click straight into a specific file and fix it with no other context, like a type of muscle memory. Now I think in prompts. I'll be cooking dinner and recall something codex said to me, then imagine myself typing a prompt back. thinking in prompts seems strictly more embarrassing than thinking in code. not sure if I'll ever go back though
Show more
0
65
1.3K
33
Forward to community
there are some really interesting rumors going around related to the distillation of open-weights models (Kimi, Qwen, Minimax, etc.) and they're very related to my PhD work The narrative [speculative]: • good distillation relies on reasoning traces, normally hidden from users • Chinese labs figured out in early 2026 how to reverse-engineer reasoning from Claude Code and Codex • they were able to collect large amounts of long-horizon data *with reasoning traces included* this way • this jailbreak led to a new wave of OSS models we've enjoyed over the past few months I'm not sure how true it is, but reasoning extractability seems like a huge uncertainty around the future of open models in particular i'm curious how much having the reasoning matters. this (plus Anthropic's messaging around distillation attacks) indicates that reasoning chains are crucial for distilling model capabilities. our research ( found something different: if you train a high-quality reasoning inverter, it's often pretty easy to reconstruct useful traces from frontier models given their outputs. figuring out how to approximate frontier model reasoning traces might turn out to be an existential problem for open weights models
Show more
0
48
1.5K
127
Forward to community
one fear around superintellingence: perhaps human AI researchers use so much compute simply because we are intellectually unable to extrapolate from results we already have very possible that smarter intelligences could learn the same or much more from tiny experiments or simulations
Show more
we didn't ever need to invent Masked Language Modeling, I don't think. it was a bit silly by construction. in most alternate universes we probably skip straight to autoregressive, sorry BERT
No opinion on math PhD but I’d still strongly recommend math major. Most of us never reason. A situation arises, we act intuitively, then get RL’ed. At best we get anxious and bounce between limited hazy alternatives. But to prove a theorem you have to reason. You don’t know what precision or thinking really is until you have to sit and struggle to prove a theorem that’s new to you. The crazier the world gets, the more this skill is useful. And the world is about to get very very crazy.
Show more
This the way: using agents to write research code you don't understand is a recipe for disaster. Use agents to obsessively poke holes in your research code and improve your personal understanding.
it's crazy to say, but through a combination of excessive paranoia and fairly standard use of coding agents, i have found bugs in - PyTorch - PyTorch/ao quantization - slime - Tinker - vLLM - FlashQLA GDN kernels ...all just in July. I don't really even know how to write kernels
Show more
0
50
1.5K
27
Forward to community
no comment on the models themselves but it's kind of funny you can just lower your prices to achieve the "pareto frontier"
pangram may well one day become the world's most important binary classifier
This is our most ambitious announcement yet. Two new models: Pangram 4 and Pangram Image. Pangram 4 completely reimagines how we approach the problem of mixed authorship, by adding a tokenwise head onto the classifier to give every token a prediction given the full document context. We also wrote a 38-page technical report detailing our experiments, methodology, and evals. Pangram Image is a completely new modality, bringing our detection expertise to AI-generated image and videos. In my early testing, it has worked shockingly well, even in strange cases like real photos of AI-generated bodega menus. Today marks a huge step in the frontier of AI detection technology. I'm so excited to finally share with you all!
Show more