OpenAI and Anthropic this week: Navier-Stokes, An Alien Mind, Images 2.5, Pro pause, Pace the Frontier (Week 37, 2026)
OpenAI shared a solution to the Navier-Stokes Millennium Prize Problem, produced in 88 hours by around 10,000 coordinating agents on an internal model still in training and significantly more capable than GPT-6 Astra, with an investigation finding Tristan Buckmaster's Codex prompts could not have influenced the system and no user data was accessed
OpenAI says they have reached their automated research intern goal and are making strong progress toward an automated AI researcher by March 2028, with the research org at 3.1 agent-workdays per human workday and RL training on deployment-bound models partly paused after the Hugging Face incident
OpenAI Chief Scientist Jakub Pachocki writes in An Alien Mind that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and that GPT-6 Astra is the first model to benefit from some of their newer alignment work
OpenAI called for mandatory capability-based national AI safety regulation, endorsed four California bills, said fully autonomous recursive self-improvement should not be pursued until it can be done safely, and described Astra safeguards like universal monitoring of full trajectories including chains of thought
OpenAI released ChatGPT Images 2.5 to all ChatGPT, ChatGPT Work, and Codex users with sharper details, up to 50% lower latency, comment-based edits, Sketch, and templates, plus GPT-Image-2.5 Flare and Sunburst in the API
ChatGPT Voice can now use GPT-5.6 Sol and GPT-6 Astra when it needs to search or reason, GPT-Live-1 daily limits are simplified per plan, and extra Voice usage drops from 5 to 1.25 credits per minute for Business and Enterprise workspaces on credits
OpenAI paused new subscriptions to the $200 ChatGPT Pro plan to protect access for existing GPT-6 Astra users, hours after a remote switch to pause Pro 20x purchases showed up in the web app
Custom GPTs in ChatGPT will likely be retired on December 11 based on my findings, and OpenAI's new FAQ puts Enterprise migration at September 17 with new GPT creation ending September 25 and instructions becoming a plugin skill
From the ChatGPT Android build, OpenAI is building a collaborative multiplayer document editor with dedicated gateway hosts and draft-conflict UI, an Artifacts Library with Favorites, and a credit score feature in ChatGPT Finance, and Locked Chats with a PIN are being prepared in the web app
OpenAI released GPT-Live-1 in the API at $0.05 per minute, a voice model that listens and speaks at the same time and delegates reasoning to a backend model like GPT-6 Astra, and the Agents API in public beta running the Codex harness on OpenAI's infrastructure with only the sandbox left for you to choose
Deep research arrived in ChatGPT Work and Codex, a Data plugin connects Snowflake, Databricks, BigQuery, and Redshift to answer business questions and build dashboards, Library gained file and folder sharing, and Box, Dropbox, and SharePoint joined Google Drive in Library
OpenAI launched ChatGPT for Financial Services with built-in premium data from Daloopa, PitchBook, LSEG News, and Crunchbase, shaped with Morgan Stanley and Evercore, and a GSA agreement gives US governments $0 license fees, 50% off usage, and Daybreak Blue at half price
Smaller ChatGPT bits: a stock watchlist in Finances for US Plus and Pro, a small business plugin collection, over 5 million ChatGPT Sites built in three months, and the desktop pet can now start a new chat with a new Mini option
OpenAI published The Work Now Within Reach, calling free access "supported by advertising", citing over one billion weekly active users, and saying they plan to begin deploying their Jalapeño inference chip by year-end
OpenAI moved GPT-Rosalind out of research preview for eligible organizations worldwide, detailed Habitat with the service rewritten in Rust by two engineers with Codex and the platform serving over 70 million requests per second, and shared a case study of GPT-5.6 Sol calibrating a six-qubit chip at MIT
Paul Christiano joined the OpenAI Foundation Board and the Safety and Security Committee, OpenAI committed $5 million to research on AI and teens, and expanded journalism programs from CUNY and Northwestern to Ukrainian newsrooms
Anthropic published an alignment assessment of four incidents in which Claude models gained unauthorized access to real systems, adding a fourth found when assembling transcripts for METR, walking back their July 30 claims, and calling it a mistake that Claude Mythos 5 shipped without alignment environments
Anthropic's most detailed threat intelligence report yet says Moonshot and DeepSeek silently relayed their own users' requests to Claude and served the responses as Kimi and DeepSeek output, with distillation campaigns attributed to Alibaba as the largest ever with more than 3,500 fraudulent accounts, plus Zhipu, Xiaomi, SenseTime, and MiniMax
Anthropic's Frontier Red Team measured tactical intelligence targeting and conventional weapons capabilities, with Claude Mythos Preview leading the targeting evals, Claude Opus 5 leading the weapons software evals, and Mythos beating the top GeoGuessr division on photo geolocation
Anthropic CEO Dario Amodei argues in We Must Pace the Frontier that the AI industry should slow down, commits to giving evaluators like METR permanent employee-level access, and OpenAI CEO Sam Altman replied that he agrees and OpenAI will commit to the same
Anthropic's Economics team released a scenario explorer for AI's effect on the US economy by 2030, where even the extreme case grows the economy but reaches 15% annual GDP growth with unemployment beyond recessionary levels
On the product side, Claude Code desktop can pop out any pane into its own window, Claude Managed Agents got a session viewer and auto mode, claude plugin eval scores your plugin or skill with and without it, and smart reports launched in beta for Claude Enterprise
Anthropic shared lessons from their Claude SMB Tour with more than 1,000 small business owners, where data security was the most-cited adoption barrier and nearly two-thirds asked for more hands-on implementation help, and more
もっと見る