Register and share your invite link to earn from video plays and referrals.

Shashwat Goel
@ShashwatGoel7
Training Decision Assistants Past work: Training AI Co-scientists, ΔBelief-RL, Measuring Long Horizon Execution
2.5K Following    4.2K Followers
I'd massively reduced podcast consumption this year bec most keep repeating the same things... But this one has the most pristine tokens on AI research I've come across. Someone should distill these into the models and we'll see progress in research taste 🤣
Show more
I went digging into OpenAI consumer training opt-out policy wordings, and I wonder if there's a loophole that allows training on hidden CoT reasoning which makes the guarantee vacuous. From OpenAI's consumer terms, opt-out apply to Input and Output (together called Content), defined as: - "Your Content. You may provide input to the Services (“Input”), and receive output from the Services based on the Input (“Output”)". ... - "Opt out. If you do not want us to use your Content to train our models, you can opt out by..." Notice, we do not receive hidden CoT from OpenAI, so its unclear if it counts as "Output" in this definition. This is important, because hidden CoT mostly contains a lightly processed version of everything useful for training, including user input and the output returned to users. Some evidence that made me particularly suspicious: other parts of the terms imply Hidden CoT is not considered Content, which then means opt-out would not apply to it: - "You are responsible for Content, including ensuring that it does not violate any applicable law or these Terms..." - "Ownership of content... (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output." Notice, since hidden CoT is not shown to us, we cannot be responsible for it, nor own it, which tells us it is not part of Output or Content, and thus exempt from training opt-out? 2. If we contrast to other similar services, e.g. OpenAI's own enterprise terms definition of Output, or Anthropic, the user receiving it is not mentioned. It is not clear whether hidden CoT is included in the opt-out in these or not either, but at least the earlier reasoning chain breaks in that "user receiving" is not explicitly mentioned, though similar statements about users owning and being responsible for Output apply. - OpenAI Services Agreement Definitions Section: "“Output” means output from the Services based on the Input." - Anthropic Consumer Terms: "Our Services may generate responses (we call these “Outputs”)" I am not a legal expert, and did not consult one, so it would be great if someone who knows better (or OpenAI) could confirm. I contacted the OpenAI data policy offer email provided 2 days back, and did not receive a clarification. Note that I am only pointing out an ambiguity in the guarantee, and not saying OpenAI definitely trains on hidden CoT. A clarification would be super useful in any case!
Show more
Sometimes papers change your world model, and I think this is one of those. The attack is super elegant, and strikes at everything Hidden CoT was supposed to protect against (a lot!) Sasha strikes again! Inspiring to see this project being cooked from across the desk :P
Show more
Prime-agent is indeed a good general harness for long-horizon tasks GLM 5.2 results on FutureSim Q2
really cool benchmark for long-horizon test-time adaptation gpt-5.5 in codex leads on FutureSim, where agents interact with a chronological replay of real-world news and are tasked with predicting future events on some Polymarket questions, gpt-5.5 even moved ahead of the human market aggregate interestingly, gemini 3.1 and opus 4.7 are missing
Show more
I have to say :) the events were discrete (but yes, noisy!) I don't think the inference from FutureSim should be LLMs suck. The environment is long horizon and relatively open-ended, GPT 5.5 still does surprisingly well. And we know RL makes them better
Show more
@ziv_ravid Continuous, high-dimensional, noisy data. LLMs totally suck at those.
Did you know you can run multi-agent experiments in FutureSim? Among many interesting behaviours, deepseek agents start updating towards the aggregate prediction on each Q, which is observable to them. Many fascinating things to explore here, like agent created+resolved qs 👾
Show more
So what happens when we make 3 copies of DeepSeek v4 Pro compete? Yes, you can run multiple agents on on FutureSim. Even when agents can only see the aggregate prediction for each question, they start converging towards it, despite being given incentives to differentiate.
Show more
Future prediction benchmarks are trivially scalable, uncheatable (under some assumptions) and impossible to saturate. They should be getting more love.
Ofc, we only do it for forecasting. But if you read the paper you can see where we are going. Every eval, from coding to writing, can follow a similar format. Temporally evolving tasks and context, agents choose what to do. This is how life the world is, so it's necc for AGI
Show more
Ofc, we only do it for forecasting. But if you read the paper you can see where we are going. Every eval, from coding to writing, can follow a similar format. Temporally evolving tasks and context, agents choose what to do. This is how life the world is, so it's necc for AGI
Show more
new forecasting benchmark: FutureSim GPT-5.5 performs the best at 25%, but Mythos, Gemini 3.1 Pro and Opus 4.7 are not included. Based on their Brier Skill Score the models don't seem to be much better than just assigning equal probabilities to all outcomes
Show more
What else have we been up to? As models get better and work over longer and longer time horizons, how do we even evaluate how well they can act and adapt? One domain we really like there is forecasting, as a hard task that test reasoning under uncertainty. We've made a benmchmark out of this, where we simulate a whole 3 month period of news, and sanboxed let models continuously read news from those days, plan, and update their forecasts. (see the animation below, just don't be fooled by its speed, this is a slice of the larger 12m token trajectory) Many more details linked below:
Show more
I've released 4 quite distinct evaluations in my PhD now: code falsification, long-horizon execution, research plan generation, and now forecasting agents. GPT models have been far ahead everytime. The wide range of tasks OpenAI post-training generalizes to is just 🤌 Never switched from codex :)
Show more
So what happens when we make 3 copies of DeepSeek v4 Pro compete? Yes, you can run multiple agents on on FutureSim. Even when agents can only see the aggregate prediction for each question, they start converging towards it, despite being given incentives to differentiate.
Show more
Continual learning is bottlenecked by realistic evaluations Introducing FutureSim, which replays real-world events in the temporal order they occurred We benchmark frontier agents at updating predictions about how our world evolves, in native harnesses like Codex, Claude Code
Show more