Register and share your invite link to earn from video plays and referrals.

Jason Wei
@_jasonwei
ai researcher @meta, past: openai, google 🧠
733 Following    112.1K Followers
Cognitive reward shapes in sports and career Sports are amazing environments to learn. When you play a sport for thousands of hours, you start to see the world through that sport. It is a simple fact—your biological neural network is being conditioned to respond to the behavior incentivized by the rules of the sport. The funny thing is that most people choose their sports for accidental reasons such as parents, geography, or school programs. People rarely think about how the particular sport you play influences how your brain thinks more generally. Going a step further, playing the right sport may even benefit your career. My two favorite sports are tennis and soccer. Tennis is one of the best sports for teaching consistency. In tennis, there are hundreds of points in a match, and each point is worth exactly one unit, regardless of whether your opponent made an unforced error or if you constructed the most beautiful point ending with a winner. Tennis is low-variance optimization—you win by reducing unforced errors, playing percentages, and grinding out small advantages. Tennis is also an individual sport, which teaches you to rely on yourself consistently. Tennis has a similar cognitive reward shape to professions like being a surgeon or a pilot. Surgery and aviation require consistency, self-accountability, and deep focus. And similar to how you can only win one point at a time in tennis no matter how spectacular it was, there is no extra credit for the best appendectomy or the smoothest SFO-JFK flight. Your craft is to provide consistency with very low tolerance for error. On the other hand, the tennis mindset transfers relatively little to entrepreneurship. Entrepreneurship is a high-variance, team game where failure is tolerated and occasional creativity gets rewarded exponentially. Minimizing unforced errors in tennis is a totally different mindset from deciding whether to make a moonshot business move that will likely fail but could potentially net a billion dollars. Obviously I am not saying that tennis players cannot be great entrepreneurs, but I do think it is a totally different cognitive reward shape. Being a forward in soccer has a much closer reward shape for entrepreneurship. What a forward in soccer learns is to create many small chances. It is a fact that most of the game, you are not scoring—even if you look at all the times that Mbappe got on the ball in one of his best games, most of those led to nothing! But all that matters is creating enough chances to score once (or a few times) and win the game. If you break down a 90-minute game for a forward, almost all the time is failure or noise, a few minutes will be leverage, and a few seconds will determine the fate of the game. I have not played soccer for thousands of hours, but I can imagine that being a lifetime forward in soccer would teach you to be comfortable with failure and asymmetric returns. In summary, I am claiming that there can be substantial value when the cognitive reward shape of your sport mirrors that of your career. I’ll admit that I’ve done some cherry-picking for illustration purposes—entrepreneurship also requires consistency and error avoidance; and goalies in soccer have reward shapes that are very different from strikers. But I think the point stands. If sports shape how we perceive risk, effort, and reward, then we should choose them wisely.
Show more
When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong enough "cognitive core", say 1B parameters, and anything else could be done with tool use, like browsing the internet or executing code. I think a lot of people were sympathetic to this argument, and indeed it is pretty hard to come up with a meaningful task that cannot be in principle achieved by a 1B model with adequate access to tools. For example, any esoteric fact that a large language model would know can be, in principle, retrieved from the internet and reasoned over by a 1B language model. However I now think this is totally wrong for one simple reason: doing tasks quickly and naturally without tool use matters a lot. The way that I internalized this reason was actually in my personal journey learning badminton this year. In badminton I am very much like a "1B cognitive core". While I can physically do every movement in a badminton shot that my coach teaches me, it requires a lot of work to mentally remember every cue and put it together. In practice I can do a shot almost perfectly, but I struggle to do it across a point and I definitely can't do it consistently in a game. This is obviously different from someone who has practiced a shot ten-thousand times and effortlessly executes it as a natural instinct. In the same way, language models knowing a fact internally, without tool calls, is meaningful. The first reason is that we obviously care about speed; you'd much rather get an answer immediately than have the model think a long time to be sure of its answer or browse the web. A second reason is that there are some things that are simply best learned via backpropagation over lots of data. If you ask about how people generally think of the Shambhala music festival, you'd rather a large language model give you an aggregate opinion based on all the data on the internet, than get a regurgitation of the first three reviews that show up in a web search. A third reason is that having to do a lot of work to find an answer is not as reliable as already knowing the answer. While this does not have to be true in theory, it is probably true in practice, at least for now. If you have to re-look up facts or redo a mathematical derivation all the time there is a higher chance of mistakes, which can compound in a long-horizon task. Once you buy that it is valuable to do things parametrically without tool use, then you must buy the argument that a 1B cognitive core is not sufficient. There is an information limit to how much knowledge can be internalized by a 1B model, and we will surely want AI to know more than that. Even 1T probably won't be enough. We will want the AI to know as much about our world as possible, we will want it to be updated with new information, and our expectations of what AI can do for us will continue to grow. In summary, tool use enables small models to do a lot more, but those who demand the highest quality intelligence will always want larger models. Bitter lesson strikes again.
Show more
0
74
1.1K
110
Forward to community
What's left for humans in a world where machine intelligence has so many advantages? I recently got a Tesla, and using full self-driving has been a wake up call to just how many advantages AI has over humans. The few times I disengaged it because I thought it was going into the wrong lane, it turned out that the car was right and I was wrong. I realized that there is no hope of me driving better than a neural net that knows every road, sees in every direction at once, and never gets tired or distracted. Given that AI has certain inherent advantages over human intelligence, what kind of moats will remain for us as humans? It's a big question. One short-term answer is that the world we live in was created for humans, and in some domains, AI has not closed the gap yet. For instance, AI still struggles to use internet user interfaces. While any computer-literate human can navigate a web page with ease, AI is still not great at making accurate clicks and drags because image embeddings are not optimized for such precision. If the internet were designed to be fed into language models instead of rendered as visual interfaces for humans, AI would obviously far exceed humans. But for now, language models still need to be retrofitted to our legacy infrastructure. Robotics is another area where we humans have a home-field advantage. Most tasks in the physical world are designed around fingers and opposable thumbs, which have been pretty hard to build into robots so far. While it is clear that machines can outperform humans in environments optimized for automation, like large-scale manufacturing lines, for now, most of the world is still built for humans. However, these capability gaps are only temporary. There will surely be a day when machines click faster than us and have superhuman general dexterity. What are the real moats that humans will have? Anything involving private knowledge that language models do not have access to feels like a solid moat to me. Romantic matchmaking and high-end real estate are two examples where inventory is often not advertised publicly and matches are made through being in the right circles. Venture capital is another example—although some research and decision making can be automated with AI, much of success hinges on understanding trends ahead of time and connecting the right people, both of which require private knowledge. While machines can and probably will have increasing access to some types of private knowledge, I think there will still be some types of private knowledge that only humans know. I do not see a path for AI to win when critical knowledge is closely guarded in human circles. A second area where humans seem to have a real moat is in entertainment and the arts, which are inherently valued for their human aspects regardless of how well machines can do them. Watching Usain Bolt sprint one-hundred meters is beautiful as an expression of the peak of human ability, even though cars can drive much faster. Watching chess at the amateur or intermediate level is more relatable and satisfying than watching two superhuman AIs play each other. The value of art comes from the creation process, which is why replicas are not as valuable as originals. These types of work feel like they will continue to have markets even as we advance towards superintelligence. More broadly, human presence is a feature that will be, by definition, challenging for AI to automate. For example, a teacher remembering your name or a parent supporting you is valuable even though AI can easily remember your name and probably give better life advice. Someone spending part of a finite life on you counts because their time runs out. As a personal anecdote, I remember the first time I worked with someone who I considered an amazing AI researcher. His advice was solid but what was more important was that I believed I could do great work with him as a collaborator and I raised my own standards. Over the past decades, the development of technology has divided us in some ways, but hopefully AI brings us closer to a world where human presence is reemphasized. Intelligence has been the defining feature of humans and it will be a big change for AI to automate that over the coming decades. In the near term, certain types of intelligence will become very cheap and automate away old jobs, but the moats I described above will not be the only places where humans can hold value. In the same way that computers took away the jobs of secretaries and manual accountants but created far more jobs via the IT industry, I believe there will be much more demand for services created by productive AI-augmented humans, perhaps for services we cannot yet imagine in today’s society. Just seeing how this story plays out will be an adventure in its own right.
Show more
On HealthBench Pro, Muse Spark 1.1 achieves similar performance with GPT-5.6 Sol (maybe slightly better) at a fraction of the cost. Affordable health superintelligence is our north star!
We benchmarked Muse Spark 1.1 and GPT-5.6 Sol on HealthBench Professional, OpenAI's benchmark of 525 real clinician tasks 🏥🩺 Muse Spark 1.1 tops our board: better overall score than GPT-5.6 Sol, statistically on par on the length-adjusted score at a fraction of the cost ($1.25/$4.25 vs $5/$30 per M tokens in/out, ~7× cheaper on output).
Show more
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!
🔥Today, we are releasing one of the first visual reasoning benchmarks for autonomous AI diagnosis in healthcare! 🚀Introducing Radiology’s Last Exam 2.0 (RadLE 2.0) from @CRASHLabAI, an uncertainty-aware benchmark for autonomous diagnosis in radiology! ✅In the last few days, the AI frontier has moved significantly. @OpenAI launched GPT-5.6 Sol. @Meta launched Muse Spark 1.1. @xAI dropped Grok 4.5. 🙌We’ve benchmarked all frontier, open-source and medical VLMs in RadLE2.0 and the leaderboard is now LIVE! 🚨 Before AI models are handed autonomy, one question matters more than any accuracy score: Do they know when to STOP and hand over to a human? ⚠️ A confident wrong diagnosis is far more dangerous than an honest “I don’t know.” Yet most models are bad at admitting the latter! 🚀 We release five RadLE 2.0 Scores: Confidence Weighted, Reliability, Accuracy, Safety and Handover Readiness and we find that models from @OpenAI @AnthropicAI @MetaAI @GoogleDeepMind @xAI @nvidia @Alibaba_Qwen @MistralAI @MiniMax_AI all score very differently as they optimize for different metrics! 🚨But most importantly, NONE of the Models have been able to reach the average human expert baseline! ⚡️A thread on what we found and which models aced our metrics! Link to the leaderboard and technical report at the end of the thread!
Show more
Excited to share what we’ve been building at Meta Superintelligence Labs! Today we’re launching Muse Spark 1.1, our strongest model yet for complex agentic workflows — delivering massive gains in agents, computer use, coding, multimodal reasoning, and multi-agent orchestration. This is also our first API release! We’d love to hear your feedback. And we’re just getting started, super proud of the team behind it, and the larger models are training right now. 🚀
Show more
Excited to share Muse Spark 1.1 with you! It’s our first model available in the API, built to excel in agentic and coding. I’m incredibly proud of what the team has achieved in such a short time, from 1.0 to 1.1. The momentum is high, and we’re still training larger models!
Show more
In addition to agents and coding, Muse Spark 1.1 is also really strong at answering health questions, a steadily growing use case for AI. On HealthBench-Pro, Muse Spark 1.1 achieves +5% better performance than Muse Spark 1.0 and beats all competitor models except Fable/Mythos. Excited for more to come
Show more
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
Beautifully written piece by @FAbnousi about how AI for health might look like in the future The current data in health is limited because it only captures episodic clinical snapshots of what happens to our bodies The revelation is that there is so much latent knowledge in looking at regular changes in our body. Could be through lab tests or even proxy metrics like wearable data With AI democratizing the ability for people to understand their own health, we're moving towards a trend of individuals gathering health data around their bodies and leveraging AI to understand themselves Bryan Johnson is extreme but a good example of this trend Personally, I started getting Function Health blood tests every six weeks instead of the recommended six months to increase fidelity on how changes in lifestyle affect my body Of course i use AI to analyze the results and adapt, and it's been great It would be cool to something in this direction happen in a big way across the world And welcome to twitter @FAbnousi!
Show more
Medicine was built on the medical record. But most of health and illness happens outside of it. AI doesn’t create that gap — it exposes it. The question now is structural: where should this data live — under what rules, and for whose benefit? With @CelinaMYongMD in @statnews:
Show more
Fun nine months! My first week i remember we had a long dinner in the cafeteria daydreaming about the cool research directions to pursue, then going to back to our desks to write a basic script to inference llama. Now we have a pretty complete stack and our first model is out 🥑
Show more
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵
Show more
Bullish, in the coming decades majority of compute will be spent on ai for science
Today, @ekindogus and I are excited to introduce @periodiclabs. Our goal is to create an AI scientist. Science works by conjecturing how the world might be, running experiments, and learning from the results. Intelligence is necessary, but not sufficient. New knowledge is created when ideas are found to be consistent with reality. And so, at Periodic, we are building AI scientists and the autonomous laboratories for them to operate. Until now, scientific AI advances have come from models trained on the internet. But despite its vastness — it’s still finite (estimates are ~10T text tokens where one English word may be 1-2 tokens). And in recent years the best frontier AI models have fully exhausted it. Researchers seek better use of this data, but as any scientist knows: though re-reading a textbook may give new insights, they eventually need to try their idea to see if it holds. Autonomous labs are central to our strategy. They provide huge amounts of high-quality data (each experiment can produce GBs of data!) that exists nowhere else. They generate valuable negative results which are seldom published. But most importantly, they give our AI scientists the tools to act. We’re starting in the physical sciences. Technological progress is limited by our ability to design the physical world. We’re starting here because experiments have high signal-to-noise and are (relatively) fast, physical simulations effectively model many systems, but more broadly, physics is a verifiable environment. AI has progressed fastest in domains with data and verifiable results - for example, in math and code. Here, nature is the RL environment. One of our goals is to discover superconductors that work at higher temperatures than today's materials. Significant advances could help us create next-generation transportation and build power grids with minimal losses. But this is just one example — if we can automate materials design, we have the potential to accelerate Moore’s Law, space travel, and nuclear fusion. We’re also working to deploy our solutions with industry. As an example, we're helping a semiconductor manufacturer that is facing issues with heat dissipation on their chips. We’re training custom agents for their engineers and researchers to make sense of their experimental data in order to iterate faster. Our founding team co-created ChatGPT, DeepMind’s GNoME, OpenAI’s Operator (now Agent), the neural attention mechanism, MatterGen; have scaled autonomous physics labs; and have contributed to some of the most important materials discoveries of the last decade. We’ve come together to scale up and reimagine how science is done. We’re fortunate to be backed by investors who share our vision, including @a16z who led our $300M round, as well as @Felicis, DST Global, NVentures (NVIDIA’s venture capital arm), @Accel and individuals including @JeffBezos , @eladgil , @ericschmidt, and @JeffDean. Their support will help us grow our team, scale our labs, and develop the first generation of AI scientists.
Show more
Old friends, new lab
After a great time at OpenAI, we (@EdwardSun0909, @_jasonwei) recently joined @Meta Superintelligence Labs. The first month has already been so much fun building from a clean slate with a truly talent-dense team! Very excited about the compute and long term focus of the new lab
Show more
New blog post about asymmetry of verification and "verifier's law": Asymmetry of verification–the idea that some tasks are much easier to verify than to solve–is becoming an important idea as we have RL that finally works generally. Great examples of asymmetry of verification are things like sudoku puzzles, writing the code for a website like instagram, and BrowseComp problems (takes ~100 websites to find the answer, but easy to verify once you have the answer). Other tasks have near-symmetry of verification, like summing two 900-digit numbers or some data processing scripts. Yet other tasks are much easier to propose feasible solutions for than to verify them (e.g., fact-checking a long essay or stating a new diet like "only eat bison"). An important thing to understand about asymmetry of verification is that you can improve the asymmetry by doing some work beforehand. For example, if you have the answer key to a math problem or if you have test cases for a Leetcode problem. This greatly increases the set of problems with desirable verification asymmetry. "Verifier's law" states that the ease of training AI to solve a task is proportional to how verifiable the task is. All tasks that are possible to solve and easy to verify will be solved by AI. The ability to train AI to solve a task is proportional to whether the task has the following properties: 1. Objective truth: everyone agrees what good solutions are 2. Fast to verify: any given solution can be verified in a few seconds 3. Scalable to verify: many solutions can be verified simultaneously 4. Low noise: verification is as tightly correlated to the solution quality as possible 5. Continuous reward: it’s easy to rank the goodness of many solutions for a single problem One obvious instantiation of verifier's law is the fact that most benchmarks proposed in AI are easy to verify and so far have been solved. Notice that virtually all popular benchmarks in the past ten years fit criteria #1-4#; benchmarks that don’t meet criteria #1-4# would struggle to become popular. Why is verifiability so important? The amount of learning in AI that occurs is maximized when the above criteria are satisfied; you can take a lot of gradient steps where each step has a lot of signal. Speed of iteration is critical—it’s the reason that progress in the digital world has been so much faster than progress in the physical world. AlphaEvolve from Google is one of the greatest examples of leveraging asymmetry of verification. It focuses on setups that fit all the above criteria, and has led to a number of advancements in mathematics and other fields. Different from what we've been doing in AI for the last two decades, it's a new paradigm in that all problems are optimized in a setting where the train set is equivalent to the test set. Asymmetry of verification is everywhere and it's exciting to consider a world of jagged intelligence where anything we can measure will be solved.
Show more
0
59
1.6K
252
Forward to community
Bryan Johnson longevity mix is the most popular drink at ragers in SF
We don’t have AI self-improves yet, and when we do it will be a game-changer. With more wisdom now compared to the GPT-4 days, it's obvious that it will not be a “fast takeoff”, but rather extremely gradual across many years, probably a decade. The first thing to know is that self-improvement, i.e., models training themselves, is not binary. Consider the scenario of GPT-5 training GPT-6, which would be incredible. Would GPT-5 suddenly go from not being able to train GPT-6 at all to training it extremely proficiently? Definitely not. The first GPT-6 training runs would probably be extremely inefficient in time and compute compared to human researchers. And only after many trials, would GPT-5 actually be able to train GPT-6 better than humans. Second, even if a model could train itself, it would not suddenly get better at all domains. There is a gradient of difficulty in how hard it is to improve oneself in various domains. For example, maybe self-improvement only works at first on domains that we already know how to easily fix in post-training, like basic hallucinations or style. Next would be math and coding, which takes more work but has established methods for improving models. And then at the extreme, you can imagine that there are some tasks that are very hard for self-improvement. For example, the ability to speak Tlingit, a native american language spoken by ~500 people. It will be very hard for the model to self-improve on speaking Tlingit as we don’t have ways of solving low resource languages like this yet except collecting more data which would take time. So because of the gradient of difficulty-of-self-improvement, it will not all happen at once. Finally, maybe this is controversial but ultimately progress in science is bottlenecked by real-world experiments. Some may believe that reading all biology papers would tell us the cure for cancer, or that reading all ML papers and mastering all of math would allow you to train GPT-10 perfectly. If this were the case, then the people who read the most papers and studied the most theory would be the best AI researchers. But what really happened is that AI (and many other fields) became dominated by ruthlessly empirical researchers, which reflects how much progress is based on real-world experiments rather than raw intelligence. So my point is, although a super smart agent might design 2x or even 5x better experiments than our best human researchers, at the end of the day they still have to wait for experiments to run, which would be an acceleration but not a fast takeoff. In summary there are many bottlenecks for progress, not just raw intelligence or a self-improvement system. AI will solve many domains but each domain has its own rate of progress. And even the highest intelligence will still require experiments in the real world. So it will be an acceleration and not a fast takeoff, thank you for reading my rant
Show more
0
84
1.4K
165
Forward to community
The most rewarding thing about working in the office on nights and weekends is not the actual work you get done, but the spontaneous conversations with other people who are always working. They’re the people who tend to do big things and will become your most successful friends
Show more
I would say that we are undoubtedly at AGI when AI can create a real, living unicorn. And no I don’t mean a $1B company you nerds, I mean a literal pink horse with a spiral horn. A paragon of scientific advancement in genetic engineering and cell programming. The stuff of childhood dreams. Dare I say it will happen in our lifetimes
Show more
The greatest contribution of human language is bootstrapping language model training