If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read.
(bookmark it)
288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection strategy as the only variable.
Correctness does not move on either agent. Equivalence testing bounds any effect to at most 10 to 15 percentage points.
A failure-mode triage explains why. Agents fail on implementation skill, feature design, pattern selection and exact wiring, rather than on repository knowledge a markdown file could supply. A manipulation probe confirms it, since the real AGENTS.md never converted a near-miss into a pass on either agent.
Borderline task difficulty is agent-specific with Spearman rho of 0.75, so single-agent studies draw tasks from different informative bands and reach opposite conclusions. That explains a lot of the contradictory prior evidence.
Paper:
Track more trending AI papers in our academy: