OpenAI says GPT-5.3-Codex lacks long-range autonomy, but admits they have no definitive metric to measure it.
Single-shot benchmarks cannot capture state drift, tool recovery, or loop degradation. Evaluating an agent that runs unsupervised for days is an infrastructure problem.