登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Shashwat Goel
@ShashwatGoel7
Training AI for Decision Making Past work: Training AI Co-scientists, ΔBelief-RL, Measuring Long Horizon Execution
参加 June 2020
2.3K フォロー中    4.1K ファン
I have to say :) the events were discrete (but yes, noisy!) I don't think the inference from FutureSim should be LLMs suck. The environment is long horizon and relatively open-ended, GPT 5.5 still does surprisingly well. And we know RL makes them better
もっと見る
@ziv_ravid Continuous, high-dimensional, noisy data. LLMs totally suck at those.