“Thinking with Images” has been one of our core bets in Perception since the earliest o-series launch. We quietly shipped o1 vision as a glimpse—and now o3 and o4-mini bring it to life with real polish. Huge shoutout to our amazing team members, especially:
-
@mckbrando, for relentlessly improving infra & ML to lay the foundation (his o3/o4-mini livestreams are the best I’ve seen)
-
@ZhangZhshuai, for pioneering our next-gen perception architecture
-
@jilin_14, for baking in the strongest perception priors
-
@bowenc0221, for initiating and showing early signs of life in thinking with images
- Jamie Kiros, for jumping in wherever work needed to get done
-
@dmed256 &
@hthu2017, for heroic infra efforts
-the Perception team, and everyone else at OpenAI who made it happen.
Multimodal is critical to OpenAI's path to AGI, and join us to push the next frontiers!
Introducing OpenAI o3 and o4-mini—our smartest and most capable models to date.
For the first time, our reasoning models can agentically use and combine every tool within ChatGPT, including web search, Python, image analysis, file interpretation, and image generation.
もっと見る