super cool work by
@eliebakouch! i noticed some comments/posts that are claiming this isn't really about generating novel research and models aren't capable of doing real research yet. imo they’re missing the most interesting signal in the traces. the agents are mainly differentiated on whether they can execute a research process in a coherent way across many hundreds of experiments. this also shows that autonomous research will be constrained as much by experimental judgment and maintaining coherent beliefs over long horizons as by generating new ideas.