登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
参加 November 2016
490 フォロー中    1.4K ファン
Thanks for trying the task. This is really cool to see! It highlights two related questions: how independently an agent can complete a task, and how human input can improve the result. To evaluate autonomy fairly, we anchor the human simulator to the original user’s actions. Your result also raises a broader question: what makes AI genuinely helpful? Task score may miss differences in simplicity and elegance. Beyond autonomy, we need to understand how well AI uses human input to produce better outcomes, and how to define and measure effective human-AI collaboration.
もっと見る
Very interesting paper. It is framed as: model intelligence is inversely proportional to # of human interventions However, I think that having a human might give a higher quality result, even if you could do it with less interventions. I used Solveit to implement & score myself on one of the tasks (very easy & reproducible paper setup, kudos for that! It was a relatively simple data anonymization task, my score was perfect (same as the AI). The resulting code was much simpler, almost half the LoC and methods. The median score of GPT 5.5 is 0.763. I would love to see more people's results! If you are interested in trying out a task, you can just use my dialog which is fully setup and just choose a different task! Share the results if you do
もっと見る