TL;DR Self-evolving agents that write their own questions and answer them can fall into "co-cheating," where the proposer and solver quietly agree on the same mistakes. Splitting source documents to evaluate across folds fixes this and lifts performance by over 8 points.
Title: False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
URL:
Points
๐ Proposer and solver share source-derived errors, letting false agreement cycle back as reward โ the paper calls this "co-cheating"
๐ Standard Dr. Zero systems show 6.1% and 8.8% false-agreement mass
โ๏ธ CrossFit splits source documents into two folds, scoring each proposer's questions with a solver trained only on the other fold
๐ CrossFit alone cuts false agreement to 3.0%/3.7%; combined with MSV it drops to 2.0%/1.7%
๐ Average downstream Cover-EM improves by 8.8 and 8.4 points over Dr. Zero
๐งฉ Multi-hop tasks see the biggest gains, averaging over 10 points
๐ฐ Compute cost rises 1.72-2.7x over baseline, though a half-budget variant still works
What stands out: without auditing the evaluator's own training history, apparent progress can be an illusion.
#
SelfEvolvingAgents# #
ReinforcementLearning#