Register and share your invite link to earn from video plays and referrals.

Jessy Lin
@realJessyLin
cofounder @EngramLab | prev PhD @Berkeley_AI
1K Following    6.3K Followers
This is a good time to share that I’m taking a sabbatical to work at Engram full time! Our blog post highlights some of the directions we’re taking, building agents that study and learn with complementary memory systems. There's lots of research still to be done, and I’m excited to help lead these efforts!
Show more
Great working with the @EngramLab team to explore model memory as a way to improve personalization and effective knowledge recall We're particularly excited about how models improve trajectories through memory. Rather than just learning to use search tools better, memory enables more directed, efficient, and effective searches that result in far higher intelligence per token. Promising for a future where agents never start from scratch!
Show more
i'm quite excited about this! • it's cool in general that models can generate their own training data and learn from it. wasn't fully clear to me even a year ago that this would work reliably. • the thing that we're trying to build seems really important and no one has built it before • some traces from our model genuinely surprise me (e.g. the attached example, where it perfectly simulates the output of a complicated bash query, despite not being trained to do this) my main feeling overall is that there is so much to do. we see signs things are starting to work: models use their memories to produce better answers, knowing things saves lots of tokens, and all of this emerges with more compute but solving this memory calibration problem, on top of learning how to generate data in ways that scale nicely with compute, is going to take some time this is just a first step 🫡
Show more
very proud to share some of our first work at Engram! These days, we have many tools for models to learn and remember: learning knowledge in weights, writing text files, practice with RL, on-policy distillation, etc. We've been thinking about what it would look like to have an agent that uses many kinds of memory natively, the same way human experts learn over time. it doesn't look like memorizing everything in weights, or grooming an elaborate system of notes, but naturally integrating all of these together to familiarize yourself in the same context day after day. We put our agent to work in Calderwood & Harkness, a 100M token synthetic law firm from our friends at Harvey, and are super excited to show a first preview of what this kind of system could look like!
Show more
Today we're publishing our first research blog, Understanding a Law Firm through Study. We're sharing a glimpse of a future where agents are trained with native memory:
Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
Show more
0
103
3.2K
463
Forward to community
I’ve been saying for a while that we need better benchmarks for continual learning/memory — I think this is a really exciting step towards this! We’ve built an entire synthetic law firm (𝐂𝐚𝐥𝐝𝐞𝐫𝐰𝐨𝐨𝐝 & 𝐇𝐚𝐫𝐤𝐧𝐞𝐬𝐬 ⚖️), modeled after real legal work – a 𝘱𝘦𝘳𝘴𝘪𝘴𝘵𝘦𝘯𝘵 work environment where agents do many tasks over the 𝘴𝘢𝘮𝘦 underlying context. Most agent benchmarks are built in this world where a model is dropped into an independent task instance and has to do its best / explore as quickly as possible (e.g. "here's a repo, fix this bug"). But we want to move towards a world where models can build on their experience. A senior engineer who has a mental "map of the codebase" can pinpoint issues much more quickly and effectively. In the same way, lawyers build up experience, learning what arguments succeed in front of regulators, playbooks for dealing with certain cases, etc. It's the 𝘥𝘪𝘧𝘧𝘦𝘳𝘦𝘯𝘵𝘪𝘢𝘵𝘪𝘰𝘯 between law firms that makes the actual practice of law interesting. Every model (and law school student) knows a lot about law from studying the textbooks, but our goal is to build models that can compound and augment the rich, internal/private knowledge that makes firm A more successful than firm B. now it's finally possible to understand and build towards that! there's a lot more work left to do, but we've been learning a lot from @ItsJulioPereyra @nikogrupen @gabepereyra to bring these models closer to real world tasks :)
Show more
when you experience a language model doing a task that used to require you at every step and now doesn't, it's a gigantic agency rupture. we tend to equate ourselves with our tasks, our outputs, and a rupture of this kind is akin to the loss of your identity and sense of self. you are no longer required, and you had no say in the matter. it's a kind of death. however, agency ruptures with AI seem to follow a predictable pattern: - Initial agency rupture. Here you see the AI doing what you used to do, and you see only the AI. In your perception of reality the AI is highlighted, and the human work around getting the AI to do the thing, is invisible. you can see this in the way we use language: "AI solved XYZ Erdos Problem" is another way to say we're in the initial rupture phase. - Seeing human scaffolding. After a little bit of experience with the model and its new capibility—programming, writing, math, etc—we start to see the edges of what's possible with the model. Here, instead of just seeing the AI doing the task, we start to see the scaffolding required by you and other humans on either end to make sure the model does its work. At this stage, the scaffolding can feel like a secondary part of the work, but it is work nonetheless. We begin to say things like, "The AI is prompting me", or "My job is just to babysit Claude." - Agency reconstruction. After a long enough time with the model at its current level of capability, we start to reconstruct a new sense of agency that centers our own role in the work being done, and makes the model's contribution mostly invisible. It is now a tool, it fast-forwards you through the boring or repetitive parts of the work, and what used to seem like scaffolding now seems like the actual work itself—the interesting part. You've had enough experience with the scaffolding to see its nuances, and how hard it is to get the model to reliably output high quality work. (As part of this, your conception of what quality work is changes—the floor is raised, and so is the ceiling.) You can tell you're at this point because you stop saying "The AI did this" and just say, "I did this" and the fact that you used AI is implied. For example, it would be weird for me to talk to someone @every and ask them if the AI built the most recent feature of a product. Of course it did! But we're so used to the capability that it becomes invisible, and we re-center ourselves and our own conception of agency. My theory is that the ability to metabolize agency ruptures and turn them into playfulness, and curiosity is a good leading indicator for whether you're a person who likes and excels in the new AI economy or not. Each time capabilities jump, you get new ruptures for existing fields touched by AI. And you get new ruptures for naive populations (like mathemeticians) who until recently hadn't been affected. if you know what the cycle looks like, it makes it much easier to ride it without freaking out. i think most people @every are naturally good at this
Show more
my definition of science: the continual process of (a) generating artifacts that are surprising or useful and (b) explaining them these days there's exponential growth in papers, and in useful artifacts, but we don't often discover good explanations so i was happy to stumble upon this extremely thorough paper about how RL surfaces useful capabilities from base models. clear experiments w great explanations are hard to find these days
Show more
engram is nothing without its people 🙏
Our founder @jxmnop recently issued some unfounded claims that got community noted. We deeply apologize for the confusion caused by his original post, the follow-up post, and the follow-up to the follow-up post. Nevertheless, we stand by his conviction in his own takes---and in strong open-source models like Inkling. Congrats to our friends at @thinkymachines !
Show more
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
0
563
14.9K
2K
Forward to community
the team is cooking,, in many ways
In this episode, @EngramLab co-founder and CEO @dan_biderman joins @allenpark to cook Mediterranean meatballs with yellow rice and talk about building AI that actually learns from you: why long context, RAG, and compaction eventually break down, how Engram compresses knowledge into cartridges and model weights, what continual learning could unlock for long-horizon agents, why token efficiency is inseparable from intelligence, how personal models could improve like Tamagotchis, and what it takes to build the research and infrastructure for millions of continuously updated AI memories. Timestamps: 0:00 Intro 0:26 Engram’s $98M Launch and Meatballs 1:45 From Naval Special Operations to AI Research 4:32 Israeli Military Culture and Founder Maturity 7:12 Why Engram Is Betting on Context and Continual Learning 9:14 Knowledge Cartridges, Compression, and Model Intuition 14:10 Trillion-Token Company Knowledge and Context Rot 18:05 Long-Context Limits, Compaction, and Neural Memory 22:20 Test-Time Training and “Destroying Prefill” 24:31 Harvey and Holistic Enterprise Queries Beyond RAG 27:02 Personal AI Models and Tamagotchi Weights 30:00 What Belongs in Weights vs. Text 32:25 Autonomous Memory and User-Specific Feedback Loops 34:20 Token Efficiency, Model Routing, and Harder Tasks 38:03 Engram’s Research Team and Product Culture 43:02 Hiring Researchers and Infrastructure Engineers 45:25 Doing More With Less 47:41 Where to Find Engram 48:19 Final Taste Test
Show more
Claude Science is incredible. I gave it some sequencing data, and in 8 hours it did a full analysis, generated figures, wrote a paper, submitted it for publication, got rejected, revised and resubmitted, got rejected again, it is now applying for positions in industry
Show more
0
103
9.1K
670
Forward to community
this is insanely cool. engram is pretty much working on a model that can improve itself through every interaction across *all* users which will likely be the single biggest capabilities jump since o1-style chain of thought reasoning
Show more
A great team working on an important problem.
I've spent my PhD figuring out how a model learns from data, digging through piles of Spanish math problems, fanfiction, and the depths of StackExchange. Engram opens a new dimension: how a model learns from you - your world, how it evolves, and the threads that tie it all together. Figuring out how to do this well is a wide-open research problem. Join us :)
Show more
finally sharing what i've been up to! left phd end of 2025 and co-founded Engram. there are a few startups in SF right making very different bets on the right way to train AI models. this is ours: people want models that learn over time, remember details, adapt and interact like a person would everyone gets a model. your model updates ~every minute. this is the world we're building. :)
Show more
0
123
1.6K
99
Forward to community
we started a company!! so, we’re tackling continual learning: what’s the learning algorithm to take arbitrary data — documents, conversations, the models’ own experience — and make better models? how do we scale compute in the same way we’ve already seen with pre-training and inference time, but scaling on the same data we see as humans, day after day with no labels, no rewards? A lot of the ingredients are out there already (rl, distillation, long-context, sparse / param-efficient architectures, etc.). our team is at the frontier of these topics, and we’re singularly focused on this. we want to understand this problem better than anyone else in the world. nobody’s solved this problem yet, but even today it’s extremely greenfield opportunity to co-develop research & useful products. in our space, how people interact with the models defines what the data distribution is - and working on this problem end-to-end, from core science to end user, gives us incredible freedom to define the problem and imagine new kinds of experiences. i expect we’ll use models that continually learn much differently than we’re using them today. it’ll feel different when the models _just know_, and build on our thinking and direction in ways we can’t even imagine. we don’t even know the queries we’re not asking, the things we would do but aren’t able to today. i’m so excited to share what we’re doing with the world in the coming months!! and the team is extremely cracked :) tackling this grand challenge and working alongside @jxmnop @EyubogluSabri @dan_biderman @MayeeChen @__howardchen @shizhehe and many others has made every day so fun. come work with us!
Show more
New paper! LLM memory keeps improving, but this makes them *worse* as user sims. If we want to build models that can, e.g., simulate realistic students to train chatbots to be better teachers, then these models need to be able to forget like humans do 📄:
Show more
Building a good benchmark for continual learning takes a lot of thought -- it's non-trivial to make it hard in the ways that matter. excited to see people working towards this
Today, we’re releasing Continual Learning Bench 1.0: the first, realistic benchmark for measuring how AI systems can improve in online settings. Benchmarks today assume models are stateless. Each example is independent, and once a system finishes a task, it moves on as if nothing happened. But deployed AI systems should learn from experience. We tested 10+ frontier systems against novel, expert-validated tasks and find there’s still plenty of headroom for learning. (1/n)
Show more
Imagine every pixel on your screen, streamed live directly from a model. No HTML, no layout engine, no code. Just exactly what you want to see. @eddiejiao_obj, @drewocarr and I built a prototype to see how this could actually work, and set out to make it real. We're calling it Flipbook. (1/5)
Show more
0
1.1K
28.9K
3.7K
Forward to community