Register and share your invite link to earn from video plays and referrals.

Yu Su
@ysu_nlp
co-founder @NeoCognition | prof. @osunlp | sloan fellow | building towards abundance of specialized intelligence
1.1K Following    15.4K Followers
Great work by the @yutori_ai team!
𝗡𝗮𝘃𝗶𝗴𝗮𝘁𝗼𝗿 𝗻𝟭.𝟱 “𝘀𝗼𝗹𝘃𝗲𝗱” 𝗢𝗻𝗹𝗶𝗻𝗲 𝗠𝗶𝗻𝗱𝟮𝗪𝗲𝗯: 𝟵𝟳.𝟯% 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗿𝗮𝘁𝗲. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group @osunlp and Careerflow Human Data Labs. All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it. But there’s a sentiment online that computer-use models aren’t progressing quickly. Not true. In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated. So what’s next? Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions — UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS). Ultimately, we’re headed to a world where computer-use models “agentify” the long-tail of the web.
Show more
a bit of a teaser for the upcoming talk @aiDotEngineer World's Fair 12:05 pm, Room 2003
Intelligence != Expertise Intelligence + Continual Learning = Expertise I will share some thoughts about continual learning and the path we @NeoCognition are on at @aiDotEngineer World’s Fair.
Show more
Intelligence != Expertise Intelligence + Continual Learning = Expertise I will share some thoughts about continual learning and the path we @NeoCognition are on at @aiDotEngineer World’s Fair.
Show more
What does the next training paradigm look like? 0:00:00 – The big research bet the labs are making 0:02:12 – Grindability is just as important as verifiability 0:06:10 – Will RLVR alone generalize? 0:08:41 – Getting the learning back to the weights 0:15:22 – Dreaming 0:17:23 – What 2027 looks like Also on YouTube, pod feed, and Substack.
Show more
OSWorld 2 is taking CUA evaluation to the next level of complexity and realism. Congrats to the team! Glad to contribute to this important effort.
Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved? Not even close. 🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. What's new: 🎯 108 real-world workflows, each ~1.6 hours ⏱️ for a skilled human ⚙️ ~318 tool calls/task vs. ~30 in OSWorld 1.0 🌍 Grounded in authentic artifacts & stateful user profiles ⚡ Captures real phenomena: dynamic environments, streaming interaction, cross-source reasoning, implicit-state inference & more 📊 Best results: Claude Opus 4.8 reaches the highest accuracy at 20.6%, while GPT-5.5 is far more token-efficient but plateaus near 13%. No one is close to solving real computer use. 🏠 Homepage: 📄 Paper: 💻 Code: 🤗 Dataset: 🧵 [1/8]
Show more
I’m excited to announce @ScaledCognition’s $100M Series A led by @vkhosla and @KhoslaVentures! I’ve spent decades in NLP, and the gap between fluent and true is exactly where AI breaks down. @roth_dan and I put together an incredible team to build a verifiable architecture focused on truth, controllability, and alignment. The result: models that are smaller, faster, less expensive, and much more reliable.
Show more
In the eyes of sloppy AI, we are all just an ID Just feeling grateful that at least I got a good number
🎉 Congrats on GPT-5.5-Cyber's progress on CyberGym! CyberGym is part of our newly launched Frontier AI Cybersecurity Observatory, our effort to provide realistic, reproducible evaluations and continuous public measurements of frontier AI systems on real-world cybersecurity tasks. In addition to CyberGym, the Observatory includes benchmarks such as ExploitGym and CyberGym-E2E, providing a comprehensive view of frontier AI cybersecurity capabilities across vulnerability discovery, exploitation, patching, and end-to-end security workflows. Really excited to see that CyberGym and the Frontier AI Cybersecurity Observatory have become important guideposts for frontier AI cybersecurity development—enabling the community to track progress, understand emerging capabilities, advance AI systems that strengthen defensive security, and identify and mitigate potential risks arising from increasingly capable offensive cyber capabilities. 📣
Show more
Some thoughts on ownership I shared with the team this morning
0
50
1.4K
127
Forward to community
Last week, we co-hosted a closed-door panel on Auto-Research & Infra with @browserbase @pk_iv @daikonland. Massive thanks to Junjie Bai (@nvidia), Yu Su (@ysu_nlp @NeoCognition), and Corby Rosset (@corby_rosset @MSFTResearch) for the incredible insights!
Show more
how do we know an agent is “continually learning”? I think most benchmarks for evaluating continual learning in agents fall short in at least one of the desiderata: 1. Task order. There should be a sequential learning process, and thus tasks should be arranged in a sequential order. 2. Task relationship. What is the relationship between earlier and later tasks? Why are experiences from earlier tasks supposed to help later tasks? What do they share in common beyond the underlying environment? 3. Metrics. If we want to measure “continual learning”, what matters is the performance difference on certain tasks before and after the task stream (or, the sequential learning process). For example, plasticity (and stability) in continual learning should be directly measured as performance difference on newer (and older) tasks before and after the task stream. In our new work (AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents), we propose a more rigorous evaluation setup for continual learning in language agents that considers all these aspects. Check it out!
Show more
a new diagnostic benchmark for continual learning that allows us to carefully examine the change in plasticity, stability, and generalization.
🚀 Introducing our latest research: AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents Continual learning for language agents has not yet been clearly defined. How should we evaluate their ability to continually learn from experience and improve themselves on complex, long-horizon tasks? - Traditional continual learning provides a useful perspective on the plasticity–stability trade-off, but its formulation does not naturally extend to non-parametric learning paradigms today. - Many recent works on agent memory still focus on retrieval and reasoning over long (static) contexts, rather than on how agents can reuse experience from complex agentic tasks. Across coding, deep research, and language understanding and reasoning tasks, we show that carefully designed task streams and metrics are essential for understanding continual learning in language agents. 📄 Paper: 🤗 Dataset:
Show more
🚨 The full program for Agentic AI Summit 2026 is now live. 📍 Aug 1–2 @ UC Berkeley 🔥 The largest Agentic AI event ever held Last year: 2,000+ in person, 40,000+ online This year: 5,000+ in person, hundreds of thousands on livestream Want to understand where Agentic AI is headed next? Join us to get the most comprehensive view of the frontier of Agentic AI, from cutting-edge research to production deployments, covering every layer of the stack: ⚡ Infrastructure ⚡ Foundation models & capabilities ⚡ Agent frameworks & platforms ⚡ Evaluation & benchmarks ⚡ Enterprise & consumer applications; agentic AI for Science, Math, Finance, Legal, Healthcare ⚡ Safety, security & governance 📣 Also excited to announce the Startup Spotlight: Building something exciting in Agentic AI? Apply to pitch directly to 5,000+ decision-makers, investors, practitioners in the room and hundreds of thousands watching worldwide. Application form in thread🧵 Deadline: July 6, 11:59pm PT The future of AI won't just be discussed here—it will be built here. #AgenticAI# #AIAgents# #ArtificialIntelligence#
Show more
Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is that really the case? Over the past many months, my group and collaborators have been building Agents' Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work. My group and collaborators previously have created many of the benchmarks the field runs on, including MMLU, MATH, CyberGym, and ExploitGym. Today, I'm excited to share Agents' Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains. With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations. The result is both impressive and sobering. Today's agents can solve a meaningful fraction of professional tasks. But when we look at the hardest tasks, the ones requiring sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance. On ALE's hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate. The age of useful agents is here. The age of truly job-ready agents is not. We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains. 🧵
Show more
0
62
977
208
Forward to community
Memory is about compression, not just storage.
Friends at Meta, @NeoCognition offers an unlimited customizable international snack bar. We have openings for anyone who is genuinely passionate about making AI agents helpful and reliable partners for humanity.
Show more
JUST IN: Meta’s CTO says morale is near “the worst it’s ever been” — leadership will offer increased snack budgets to lift spirits.
New @NoPriorsPod: @LipBuTan1, CEO @intel and legendary semis investor I've been lucky to partner with topics: - Why take the CEO job - His ten-year vision - AI for chip design - Making chips in the US - How to invest in semis - What investors misunderstand about Intel
Show more
🚩 AI-assisted scientific writing faces increasing scrutiny Can human or LLM editing fix it? 👇 240k edits on scientific abstracts expose LLM writing flaws, and show that human editing fails to fix them cc @hsanchaita, @leadoeun27, @shocheen, @mbodhisattwa 🧵
Show more
I'd also like to share a few thoughts from the past six months of building QUEST that didn’t make it into the paper, but might still be useful to the community. > Mid-training is surprisingly powerful, especially for smaller models. In our early experiments, we applied mid-training to smaller models like Qwen3-8B and saw surprisingly large gains (~3pp over pure SFT) with only around 100K tasks. The effect became weaker as model size increased, but for smaller Deep Research agents, carefully designed mid-training tasks can still go a long way. How should those tasks be designed? That’s a much longer story. If you’re interested, check out the “Unsuccessful Attempts” section in our paper. Honestly, it’s probably my favorite section in the whole paper. > For agentic RL, infrastructure matters way more. This might sound like a cliche in 2026, but I still want to yell it loudly. We spent almost two months just making our RL pipeline stable. Agentic RL is essentially a giant chain of dependencies. The judge model needs to work. The search and retrieval services need to work. The cache needs to work. And all of them need to keep working for days. A tiny bug in any part of the pipeline can ruin an entire training run. We end up spending a surprising amount of time designing fallback mechanisms for everything. Not because it’s elegant, but because otherwise your training will randomly die at 3 AM after running for days. The secret isn’t fancy RL tricks. The secret is making sure your pipeline survives when things inevitably break. > Session-level training is probably the reason we can afford long-horizon training. Everyone working on Deep Research agents knows the pain: context is expensive. Compared to traditional reasoning tasks or short-horizon agents, long-horizon Deep Research training burns through GPU resources much faster. For most academic groups, GPUs are still the biggest bottleneck. In QUEST, we use session-level training and limit each training instance to 32K tokens. It’s not the most glamorous idea in the paper, but it helps make large-scale long-horizon training much more practical. > Cache saved us a ridiculous amount in API spend. We spent a lot of effort building our caching system. At first, it sounded like an engineering optimization. Later, it became a necessity. Many websites visited during data synthesis show up again during RL training. So as the search tool. Without caching, you’re effectively paying for the same information over and over again. The funny part is that cache becomes even more valuable when your experiments fail. RL runs crash. Services fail. You restart things. But cached results stay. One funny observation: the more failed runs you have, the higher your cache hit rate gets. It’s a slightly sad but comforting story: every failed run leaves behind a few more cache entries for the next one. By our final training run, the cache hit rate had reached ~40%, which translated into a significant reduction in API costs. How painful it was to build. How rewarding it was to finish. We hope you’re as excited to meet QUEST as we are to finally share it.
Show more
Excited to share our lab’s latest work on fully open deep research agents, led by @jianxie_! For anyone doing agentic post-training, it’s a gold mine: a full recipe (mid-training → SFT → RL), a rubric-tree data synthesis pipeline, and an open model family from 2B to 35B.
Show more
🚀Excited to announce QUEST today! Benchmark numbers will fade, but the know-how behind them endures. With QUEST: See how we synthesize diverse Deep Research tasks through a unified framework at scale. See how we train Deep Research agents through three stages: Mid-Training → SFT → RL. See how our caching system dramatically reduces API costs during training. See how our context management enables unbounded deep research. We’ve released everything we can!
Show more