Register and share your invite link to earn from video plays and referrals.

Jiahui Yu
@jiahuiyu
Startup. Previously co-founded TBD @Meta Superintelligence Labs; led Perception @OpenAI; co-led Gemini Multimodal @GoogleDeepMind.
1.3K Following    28.3K Followers
I’m leaving Meta to start a new company. Building the TBD Lab alongside Mark @finkd and Alex @alexandr_wang has been deeply inspiring and fulfilling. I’m proud of what our multimodal team accomplished across Muse Spark, Voice Mode, Muse Image, and Muse Video, and even prouder of the team that made it possible. Over time, I’ve felt increasingly drawn to a problem that will matter deeply to humanity’s future, yet remains largely underexplored. It now has my full attention. More to share as the work takes shape.
Show more
0
275
4K
119
Forward to community
"The AI Future Is for Everyone." Couldn't agree more with @finkd. One piece I personally feel strongly about: Frontier AI needs to be for everyone, too. How many people outside tech/eng actually know what to do with the top models today—or even feel the difference when a new one drops? Closing that gap is where the real future lies.
Show more
People consistently underestimate how good Muse Spark is In perceptual tasks, it beats Claude Fable in any task I've thrown at it eg (Claude Fable incorrectly answered 6 when a friend tried)
Show more
#1# on video-to-web! Recording a video is one of the easiest ways to turn an idea into code. Thanks to the @DesignArena team for the benchmark!
BREAKING: Muse Spark 1.1 takes 1st on our new Video to Website leaderboard with an Elo of 1250. Video inputs capture richer context than static images - including interactions, transitions, and responsive behavior - challenging models to reproduce the full experience. Only six labs currently support native video input, with @Meta debuting 14 Elo points ahead of @Kimi_Moonshot and 32 ahead of @GoogleDeepMind. Video input is quickly becoming one of our users’ most requested capabilities. Note: OpenAI & Anthropic currently do not yet support native video input in the API. Congratulations to the @AIatMeta team!
Show more
pleasantly surprised!
Holyyy SHiiiitttt 🤯 meta muse spark 1.1 passed this image test that even fable failed < this vision is on par of google models > @alexandr_wang what did you guys cooked , and thanks to let us see reasoning , best release for the price @AIatMeta is back in the game
Show more
Muse Spark 1.1 is here! Available today via API and on Of all the agentic use cases, I’m most excited about this one: going from a single video shot straight to task completion. 🚀 We’ve also polished and launched Computer Use and Video-Audio Perception for Muse Spark 1.1. Read more:
Show more
Muse Spark voice mode! Available in Meta AI today.
Today we’re introducing Meta AI Voice Conversations powered by Muse Spark that let you talk naturally to Meta AI (interrupt, switch topics, or swap languages), and as you talk, Meta AI can generate images and pull up recommendations from Reels, maps, and more. We’re also bringing live AI to the app, so you can point your camera at the world and ask about what you’re seeing in real time.
Show more
Fun visually-grounded 3d visualization with Muse Spark - by asking the model to generate a 3d visualization of how much each component cost on a xbox 360 motherboard (and put the bar on the component).
Show more
😂 It’s been a lot of fun working with @lindugong — come join us!
when msl first started we had 4 researchers, 1 gpu, and @jiahuiyu was wearing a facemask all the time from going viral. we’re so back!!
Happy to share Muse Spark, a natively multimodal reasoning model w/ tool-use, visual chain of thought, and multi-agent orchestration! It’s been a fulfilling journey not just building the model, but the team and culture behind it. Now live in product.
Show more
this has been one of my helicopter moments too:
0
168
2.1K
160
Forward to community
> a master's thesis worth of work lol clock reading isn't always accurate yet, but you may find these clocks from the internet interesting in one of my tests, the model reads the clock back and forth ~40 times before reaching a conclusion - scaling with test-compute for real... i'm fairly confident our models will be way more concise next time!
Show more
It took o3 practically a master's thesis worth of work, but *did* correctly read my wrist watch this afternoon!
> see how it solves a maze lol I tried a 200x200 maze and it worked, but it looked too crowded for the blog post—so we used a 25×25 maze instead.
We launched o3 and o4-mini today! Reasoning models are so much more powerful once they learn how to use tools end-to-end. Some of the biggest lifts are coming in multimodal domains like visual perception (see how it solves a maze in our blogpost: 🤯)
Show more
“Thinking with Images” has been one of our core bets in Perception since the earliest o-series launch. We quietly shipped o1 vision as a glimpse—and now o3 and o4-mini bring it to life with real polish. Huge shoutout to our amazing team members, especially: - @mckbrando, for relentlessly improving infra & ML to lay the foundation (his o3/o4-mini livestreams are the best I’ve seen) - @ZhangZhshuai, for pioneering our next-gen perception architecture - @jilin_14, for baking in the strongest perception priors - @bowenc0221, for initiating and showing early signs of life in thinking with images - Jamie Kiros, for jumping in wherever work needed to get done - @dmed256 & @hthu2017, for heroic infra efforts -the Perception team, and everyone else at OpenAI who made it happen. Multimodal is critical to OpenAI's path to AGI, and join us to push the next frontiers!
Show more
Introducing OpenAI o3 and o4-mini—our smartest and most capable models to date. For the first time, our reasoning models can agentically use and combine every tool within ChatGPT, including web search, Python, image analysis, file interpretation, and image generation.
Show more
0
64
1.5K
157
Forward to community
Yesterday was crazy so i missed my post but damn one more "mini" coming soon!
Yesterday was crazy so I missed my post but damn the 4.1s awesome. @michpokrass @jiahuiyu @johnohallman @suvansh @StrongDuality and so many others did a phenomenal job. We’re truly horrible at naming but the secret trick is that the models with mini in their name are 🔥
Show more
o1 and o3-mini now support file & image uploads. more is brewing - stay tuned!
Two updates you'll like— 📁 OpenAI o1 and o3-mini now support both file & image uploads in ChatGPT ⬆️ We raised o3-mini-high limits by 7x for Plus users to up to 50 per day
Congrats on the launch @shgusdngogo ! Your progress and resilience are truly inspiring. Always glad to empower Operator with better Perception!
We're excited to announce Operator: our research preview powered by Computer-Using Agent. It can use its own browser by perceiving screens from pixels and taking actions using keyboard and mouse.