Register and share your invite link to earn from video plays and referrals.

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
Joined November 2016
490 Following    1.4K Followers
Thanks for trying the task. This is really cool to see! It highlights two related questions: how independently an agent can complete a task, and how human input can improve the result. To evaluate autonomy fairly, we anchor the human simulator to the original user’s actions. Your result also raises a broader question: what makes AI genuinely helpful? Task score may miss differences in simplicity and elegance. Beyond autonomy, we need to understand how well AI uses human input to produce better outcomes, and how to define and measure effective human-AI collaboration.
Show more
Very interesting paper. It is framed as: model intelligence is inversely proportional to # of human interventions However, I think that having a human might give a higher quality result, even if you could do it with less interventions. I used Solveit to implement & score myself on one of the tasks (very easy & reproducible paper setup, kudos for that! It was a relatively simple data anonymization task, my score was perfect (same as the AI). The resulting code was much simpler, almost half the LoC and methods. The median score of GPT 5.5 is 0.763. I would love to see more people's results! If you are interested in trying out a task, you can just use my dialog which is fully setup and just choose a different task! Share the results if you do
Show more