Register and share your invite link to earn from video plays and referrals.

Parth Asawa
@pgasawa
CS PhD student @Berkeley_EECS
381 Following    2.4K Followers
Introducing Exo 🦋 – an open-source agent that can rewrite every part of itself. It learns new environments, adapts policies, even optimize costs (and play games!!) We've been hacking away at this project with @martin_casado from @a16z and @ankrgyl from @braintrust and friends.
Show more
We expect agents to act like senior engineers, but most benchmarks still evaluate them like interns. Excited to introduce Senior SWE-Bench, an open-source and @harborframework-native benchmark that assesses agents as senior engineers on long-horizon tasks with realistically under-specified instructions. We expect agents to build real features going on just a quick Slack message, nothing like the super technical instructions most benchmarks provide. Senior SWE-Bench fixes that. Claude Opus 4.8 is the current leader at 24% high quality solves, but it took 117K tokens on average to get there. Claude Sonnet 5 looked like it was going to swoop in for the top spot, but we found it cheated on 26% of trials.
Show more
I'll be at ICML July 7th-10th, hit me up if you want to chat about continual learning, AI policy, etc! I’m giving an invited talk on evaluating continual learning at the CATS workshop on July 10th at 8AM KST and @aczhu1326 and I are presenting a poster on Advisor Models on July 8th at 10:30AM KST (Hall A #2107#).
Show more
New episode of Benchtalks is live. @pgasawa (Berkeley PhD, advised by @matei_zaharia + @profjoeyg) built Continual Learning Bench. The field didn't have a shared definition of continual learning, let alone a benchmark. Parth fixed that. Full episode with @vincentsunnchen:
Show more
Super fun chatting with @vincentsunnchen about all things continual learning including evaluation, parametric methods, open science, and more!
capability != learning new benchtalks with @pgasawa on continual learning, where we discuss teaching models to learn from experience, measuring learning ability, the bet on parametric models, and more 01:06 What is continual learning? 04:10 Why capability and learning are different 06:13 Why build a benchmark? 08:07 Continual Learning Bench launch and reception 09:13 Anthropic's Fable release and Continual Learning Bench 11:02 How to design tasks for continual learning 18:41 The gain metric 24:01 What good looks like on the leaderboard 29:13 Failure modes: why models can't update their beliefs 31:12 Parametric systems and future architectures 34:30 Open science and AI safety 45:42 Lightning round 49:08 How to contribute to Continual Learning Bench
Show more
Always fun to chat with @andykonwinski, see 41:33 for some of my brief takes on evaluating continual learning and academia / open science.
At @CAISconf last month, @andykonwinski sat down with researchers on the conference floor -- @matei_zaharia @istoica05 @lateinteraction @dawnsongtweets @gneubig @pgasawa @JonSaadFalcon @heathercmiller @ryanmart3n @alexgshaw @profjoeyg @swyx and Ioannis Ioannidis -- to talk open vs. closed, agent evals, compound AI systems, and where the frontier is headed 👀
Show more
Benchtalks Ep. 3 with @pgasawa (Continual Learning Bench): coming soon with @vincentsunnchen 👀
Databricks and Perplexity co-founder @andykonwinski has a phrase for the AI industry's current trajectory: "feudalism with better branding." With @LaudeInstitute, he's trying to fight AI's concentration of power by incentivizing top AI researchers to do their work out in the open. I've known Andy for over a year now and find him incredibly thoughtful on how AI is impacting society. He's on the podcast this week:
Show more
Open-science is the only thing that really needs to prevail. Good post. If only there was someone I knew building an institution like this.
What is the future of AI research? Maybe the highest leverage thing we can do isn't to train the next model or write another paper. It’s time to re-think the fundamental process by which science, innovation, and regulation happens, who participates, and what we want the future to look like. (1/2)
Show more
A good essay by @pgasawa and @profjoeyg on a more nuanced view of AI advances.
We need more rising star researchers like Parth talking about how the technology we are building will impact society. I also happen to agree with him.
This week made something clear: it’s time to stop treating concentration of power in AI as a solution rather than a risk. Safety and centralized control are not the same thing, so let's stop talking about them like they are. Yet scaling laws are real. The threat of AI cyber-hackers disrupting the global economy is real. The threat of somebody using AI to create a biological weapon is real. AI safety is a species-level concern and we need the best minds in our species from across institutions working on solutions, not locked outside the closed doors of frontier AI development. We need a shared research commons at the intersection of industry, academia, and the public good—an open research ecosystem with access to billions of $ in compute, SoTA models, and strict protocols for dissemination that ensure the most impactful discoveries in the history of our species are shared in a safe way that benefits all of humanity. Not blind open source ideology. Not closed access as safety theater. An open frontier, shaped by many and accountable to all.
Show more
Awesome to see folks using tasks from Continual Learning Bench to evaluate what aspects of continual learning can look like with models like Mythos/Fable!! I'm eager to push the frontier of training methods and evals here -- reach out if you want to chat about all things CL :')
Show more
// Continual Learning Bench // One of the research areas with lots of investments is continual learning. While there are many efforts, there is very little progress in measuring it. So the big question is, do dedicated memory systems actually make agents learn from experience? Continual Learning Bench says not yet. Across six expert-validated domains with shared learnable structure, naive in-context learning outperforms systems purpose-built for memory management. CL-Bench introduces a gain metric that isolates genuine learning from prior capability, then shows agents frequently overfit to immediate observations or fail to reuse knowledge across instances. If a plain ICL baseline beats your memory architecture, the architecture is adding overhead rather than learning. Paper: Learn to build effective AI agents in our academy:
Show more
We release Recon — a new approach to reasoning synthesis for user modeling. The key insight: post-hoc rationalization ≠ reasoning. We propose using action reconstruction as a scoring criterion for synthesized reasoning traces, yielding more causally faithful reasoning and improved downstream action prediction across user modeling tasks. Paper and project page in 🧵
Show more
Really fun conversation with @swyx and @jacobeffron. AI in Healthcare can have tremendous impact, yet I find many may not be aware on technical problems in the field. I can tell you, as we cover in this podcast, there are so many frontier technical opportunities in a large and greenfield industry!
Show more
Today, we’re releasing Continual Learning Bench 1.0: the first, realistic benchmark for measuring how AI systems can improve in online settings. Benchmarks today assume models are stateless. Each example is independent, and once a system finishes a task, it moves on as if nothing happened. But deployed AI systems should learn from experience. We tested 10+ frontier systems against novel, expert-validated tasks and find there’s still plenty of headroom for learning. (1/n)
Show more
0
42
1.2K
167
Forward to community