Pretty interesting rethinking paper on "discouraging" self-evolving loops.
So the background is that most current self-evolving loops run their search directly on the test set. It kind of makes sense as harness search needs accurate, verifiable feedback to make grounded edits.
However, that quietly turns self-evolving loops into a form of test-time scaling, which is exactly what this paper argues.
Specifically, the paper points out that a loop that repeatedly evaluates and revises candidates against task feedback, then reports on those same tasks, is logically a test-time search procedure. So its gains should be measured against test-time scaling under matched feedback and inference budgets — otherwise you can't tell whether it discovered a better harness or just spent more compute.
With that in mind, the paper runs four methods under the same budget:
1. parallel sampling — fixed harness, k independent trajectories per task, with a self-judge or unit tests picking the final answer
2. sequential refinement — fixed harness, k retries in a row. Each round summarizes the previous attempt into context and tries again (essentially prompt refinement)
3. harness evolution — the standard self-evolving loop. One shared harness, revised each round from feedback pooled across all tasks
4. harness scaling — the per-instance counterpart. Each task evolves its own harness
The results are very interesting.
Harness evolution doesn't beat plain parallel sampling, and without verifiable feedback it can even fall below single-attempt direct sampling with the initial harness.
Its gains also show up at pass
@5 but barely at pass
@1, which implies that the improvement comes from taking multiple attempts, not from the harness getting better.
And on a disjoint search/eval split the evolved harness transfers almost nothing to held-out tasks, which means the edits memorize task-specific fixes rather than distill reusable strategies.
There is one caveat, which is that the "unified budget" only counts inference on the tasks, not the compute spent generating harnesses. But this flaw kind of favors harness evolution, and it already loses.
So in my opinion this really shows that existing self-evolving loops might just be a different way of applying test-time scaling, rather than some new intelligence discovery.
And we should focus more on making generalizable self-evolving loops work!