Follow-up: Agent Readiness was most definitely worth the run.
The new tests identified a few bugs that were materially impacting a data ingestion cron job, and things were failing silently.
Multiple runs on the other harness - with different model permutations - didn't catch these.
Harnesses matter. Agent runtime matters. Models can be flexed differently.
I'll report back on the /missions output when I get time to review, approve, build and deploy!