I have to say :) the events were discrete (but yes, noisy!)
I don't think the inference from FutureSim should be LLMs suck.
The environment is long horizon and relatively open-ended, GPT 5.5 still does surprisingly well.
And we know RL makes them better