I've released 4 quite distinct evaluations in my PhD now: code falsification, long-horizon execution, research plan generation, and now forecasting agents.
GPT models have been far ahead everytime. The wide range of tasks OpenAI post-training generalizes to is just ๐ค
Never switched from codex :)