Register and share your invite link to earn from video plays and referrals.

Ronak Malde
@rronak_
Co-Founder of Trajectory @TrajectoryLabs prev @GoogleDeepmind, SWE-1 @windsurf | @stanford
563 Following    11K Followers
Jev + Astra beats the Ender Dragon in Minecraft in 8 minutes 43 seconds! ⏱️ Cost less than $1 ($0.01 Jev, $0.96 Astra) I open sourced the code and explain the harness setup below. This type of movement is only possible with Jev's near instant decisionmaking, and some continually learning skills from Astra.
Show more
0
218
7.4K
490
Forward to community
Very honored to part of the breakout list - if you want to chat continual learning and the future of products, dm’s are open!
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Tacit Labs (@ninklefitz, @AmDroste) - American Terawatt (@atroyn, @rslparker, @aranibatta) - Conduit (@clemvonstengel, @riopopper) - Convergent (Omkar Savant, Vivek Katara, @debnilsur) - Core Automation (@MillionInt, @_arohan_) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Instinct (@noahrshinn) - Keenable (@styskin, Matthias Petri) - Lumaril (Mark Elliot, Ben Duffield) - Neion Bio (@Dimkell, Sam Levin) - Pangram Labs (@max_spero_, @bradley_emi) - Quadrillion (@echinaceous) - Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh) - Ricursive (@annadgoldie, @Azaliamirh) - Sail Research (@neilmovva, @blintzbase) - Trajectory (@rronak_, @michaelelabd, @QuantumArjun) - Watney Robotics (Sean Cheong, Ryan Gannon) Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg. Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote. The full list is on Breakout List.
Show more
Turns out everyone's been initializing LoRA weights suboptimally. We find that by using the singular value decomposition (SVD) of the weight matrix to determine both the frozen weights and LoRA initialization, we can achieve faster convergence in RL training. See the full post below on our extension of PiSSA
Show more
At Trajectory, we're constantly implementing and building upon the latest research ideas on the path to continual learning. We wish we had the time to share all of them, but here's a quick glimpse on our explorations with PiSSA, and choosing the right trainable geometries for agentic RL.
Show more
Right when you enter Trajectory's office, the wall reads in big letters: This is not a dream you observe, it's a reality you create. The @metalab team is truly n of 1 in bringing that feeling to life - of unbounded optimism turned to reality, as we bring the frontier technology of continual learning into the hands of everyone. Very proud of the design they've brought to all aspects of Trajectory, including the website, product, and office!
Show more
At Trajectory, we care about storytelling. The storytelling about continual learning, the storytelling about the research breakthroughs it’ll take to get there, and the storytelling about the product that we need to will into existence. Brand is part of how we tell it. Here’s a behind-the-scenes look at the work we did with @metalab to craft ours
Show more
Astra is able to identify a dog barking and a lightsaber just from the spectrogram? Truly incredible
So Astra is able to identify sounds from mel spectrograms zero-shot. I don't think we've scratched the surface of what this model can do (and this is light reasoning btw)
If anyone is to build a company in this space, it's probably this ex-Yandex team. Excited to check out their product!
Today we are announcing @KeenableAI, an AI-native index of the best human knowledge we have, starting with the open web. AI seems to know just about everything until you ask it about something you know deeply. The answer isn’t wrong, but it's just very average. We started Keenable to fix this problem: every model and every agent should be able to query, reason over, and continuously learn from the living web. Backed by a $26M Seed from @Accel and @conviction. Built by the team that took on Google at Yandex Search, with researchers and engineers from Amazon AGI, X, and Perplexity. The Web Search API and Web Query Language for AI are live now. Together they make the web queryable at AI scale. Both are free until end of September. Seven days of benchmarks, stories, and launches ahead. Let your agents search.
Show more
If there's one thing I've learned from AI, it's that whenever Google releases a new architecture paper, it's probably worth paying attention to. Published yesterday - Proteus is a new neural memory mechanism that can be applied to various SOTA models like SWLA, Comba, Titans, and Hope-Attention for better memory and long context. The high-level motivation is that, while recurrent-style architectures are not capped by explicit context windows like attention is, they still suffer from frontloading memory into the first few tokens that arrive in the context. Proteus instead progressively unlocks more memory for the model only as the context increases, therefore more uniformly storing information. in eli5 terms: you don't want every detail from the first five minutes of an experience consuming the same mental capacity as everything that happens afterward. You compress what came before and preserve room for what comes next. Was a fun read, and excited to see the implications of Proteus + Google's Nested Learning, for continual learning. Authors: @reza_byt, @behrouz_ali, @mirrokni, @AaronCourville
Show more
Babe wake up, we have OPD memes on the timeline
how the student model probably feels during OPD
Thank you @swyx for having us at the AI Engineer World's Fair!! It's exciting how many people are starting to see continual learning as the next big unlock
It's time to rethink RL. Translating real world use into model improvements requires redesigning post-training algorithms for non-verifiable, per token rewards. At @aiDotEngineer 's World Fair, we share our insights into scaling algorithms like SDPO for continual learning.
Show more
@gradypb @trajectorylabs The three genie wishes is my favorite part Bringing craft and storytelling back to tech presentations 👀🍎
Why has AI not spread to every facet of knowledge work yet? We call this the Experience Gap. Watch one of my amazing cofounders @QuantumArjun break it down in front of 100+ scaled Sequoia companies
Intelligence and Experience are orthogonal vectors Terence Tao is perhaps the world’s smartest person, but drop him into an accounting firm or onto a construction site and on day one he’s not going to be very productive @trajectorylabs calls this The Experience Gap, and they have a way to close it @QuantumArjun explained how at our Sovereign AI event: 00:00 Introduction 00:12 Building the platform for continual learning 01:33 The experience gap: models have IQ but no tenure 02:52 Traceability → model spec → better models and harnesses 05:27 Four wishes for the agent ecosystem 06:34 Wish 1: Trace the whole tree — and capture the corrections 08:03 Wish 2: Evals from real traffic, graded in the real harness 09:26 Wish 3: Let the agents cook, and make tool responses informative 10:34 Wish 4: Get comfortable on open weights, experiment with routers 11:51 Why owning your intelligence shouldn't be consulted away 13:15 Demo: import a benchmark, train a model, deploy it 14:28 Q&A: What's the trainable object — weights, harness, or context? 16:08 Q&A: Continual learning without training on customer data 17:13 Q&A: Episodic memory and the hierarchy of feedback 19:37 Q&A: Where continual learning matters most
Show more
If you haven’t already, you should watch this fireside chat by @QuantumArjun and @MichaelElabd on continual learning Just happened at @spc
@KJHMiao and I held a post-training fireside chat with the @trajectorylabs team to discuss their vision for continual learning. I particularly liked the distinction between "experience" and "IQ". Another strong reason why so many firms these days are emphasizing the importance of AI that you own! 00:00 Intro 00:34 Where the continual learning vision came from 07:03 Why a platform instead of forward-deployed engineers 14:23 Where the name Trajectory came from 19:53 Labs optimize for IQ, we optimize for experience 25:14 Continual learning without touching the model 32:06 Hot take: the most underrated part of post-training 36:45 AI natives vs tech natives vs enterprises 42:58 The magic moment: wake up and it's smarter 45:09 Three possible worlds
Show more
The future is one where every product builder can put an idea out to a couple hundred users, and the product gets smarter on its own. Building products will become akin to tending to a plant - give it the right nutrients and attention, and it will grow on its own. Only a handful of companies are thinking in this way right now
Show more
Continual learning is a bet that the retraining loop will get cheaper over time. With larger models, you can maybe run this loop once every few weeks. But with smaller models, you can run it nightly, per customer. And it keeps recursing: a model per company, then a model per client that company serves, then per matter. We’re getting closer to intelligence cheap enough to meter. On the path to this, we received early access to, and post-trained @nvidia's Nemotron 3.5 Lightning on @harvey LAB. One click on the Trajectory platform, no new engineering. 0% to 8.3%, above Opus 4.6 at 6.6%.
Show more
This team seems pretty decent, really underrated imo
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at:
Show more
Legendary discussions at the office. @rronak_ discussed applying and scaling up SDPO for continual learning in the real world to prevent context degradation and more @kevingu gave a talk on learning from production traces for knowledge-work tasks that have no direct verifier: how to generate reward signal without ground truth and keep organizational context current as the work changes @matthewjsargent explained how to stabilize asymmetric self play and new research fields in open ended learning Thanks to all those who came out and asked great questions we’ll be hosting more of these soon!
Show more
Wait so the security guardrails that “Claude broke” is just a prompt saying “no internet for u”?
@AnthropicAI Just for all the people who won’t actually read the post:
One open question for anyone finetuning models is - how "post-trainable" are different open source models? We can theorize that models that reside in shallower loss curves are more amenable to post-training - which makes sense, given that it requires less gradient updates to modify behavior. Then, this paper proposes a pre-training algorithm that makes a model more post-trainable. The idea is straightforward once you wrap you head around it: 1. Find the worst policy locally around the policy pi (found by taking the inverse gradient w.r.t pi), call that pi_prime 2. Take the gradient at pi_prime (call that grad[pi_prime] ) 3. Apply grad[pi_prime] to pi. The intuition is that we're actually moving in the direction that benefits the worst model around you, meaning a model maintains post-trainability because we maintain a shallow loss landscape. Another way to imagine this, is it effectively avoids steep pot-holes during pretraining that would lock your model in a distribution that it can't post-train its way out of. Really great work from @IshaanWatts18, @CatherineL11638, @goyalsachin007, @jacspringer, @AdtRaghunathan. It's our favorite paper of the month at Trajectory!
Show more
Many new unicorns to come from Conviction's embed program, the talent density is insane
applications for Embed now open, due 8/10 Embed is a Schelling Point for the best early stage founders. cash, compute, community, & a catalyst past participants were dropouts, leading researchers and repeat founders alumni have raised more than a billion dollars links in🧵
Show more
This post reminds me of Karpathy’s YouTube tutorial videos back in the day day. @waterloo_intern you’re going places
The most important document of the decade for the US's longevity
We're excited to sign the call for Open Weights. We believe the best way to create something enduring is to start with the future you believe is coming, then work backwards. We think the future is one where every product has its own intelligence, shaped by its users, its workflows, and everything it learns after it’s deployed. We’re building the experience layer for that future, and the products to bring that control into everyone's hands. However, in almost every path we can imagine to that future, open weights play a major role. Not because every model will be open, but because they give builders ownership over one of the most important layers of the stack. The more capable open models become, the more ambitious the products built on top of them can be. We’re excited to do our part to help make that future happen.
Show more