登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

λux
@novasarc01
tensor shepherd in a non-euclidean pasture | grazing on cuda cores
参加 July 2024
3.5K フォロー中    22.5K ファン
absolutely loving this. for the past few months i’ve also been exploring synthetic and procedural environment generation for agents (though at a much smaller and earlier-stage scale). the projects i’ve been working on (synthetic workspace gym and rlvr gym) include a small set of environment families such as python script repair, pipeline repair, tabular tasks, scheduling, graph planning and retrieval-style workspace tasks, with workspace artifact logging, hidden evaluators, trajectory traces and final diffs. the motivation was mostly curiosity...prompts alone feel too weak as a substrate for studying long-horizon agents. what we really need are executable worlds where agents can inspect state, take actions, receive grounded feedback and fail in ways that are actually analyzable. that is why i’m especially excited by what the prime intellect team has done here. we need thousands of diverse, verifiable, stateful tasks across domains not just static prompt datasets or a handful of handcrafted benchmarks. imo scaling this lets us study curriculum, tool use, verifier design, task difficulty, trajectory quality, reward hacking and long-horizon generalization in a much more systematic way. my own versions are still early and small but it’s really exciting to see PI push this direction seriously: evolving environments, calibrating difficulty and generating broad task distributions across thousands of domains. also i really loved the clean and neat blog post!
もっと見る