A pipeline with 100% delivery rate, 100% schema validity, and zero retries or errors... yet re-running the exact same request flips the verdict. This paper reports that shocking negative result, fully preregistered with a complete audit trail.
Title: Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
URL:
๐ง Highlight 1: Clean engineering does not mean reliable measurement
Across 3,312 planned calls, response rate and schema validity both hit 100% with zero retries or errors, yet the repeat-ranking agreement (median Spearman) was only 0.400 โ clearly failing the preregistered stability gate (threshold 0.90).
๐ฒ Highlight 2: Byte-identical inputs still drift
Replaying the exact same request 24 hours later gave only a 0.780 exact-ranking agreement rate. Interestingly, comparisons across windows on the same day scored a nearly identical 0.805 median, revealing this isn't "next-day drift" but immediate platform-level nondeterminism.
๐ Highlight 3: More samples don't fix it
Scaling observer calls from 8 to 500 (748,000 calls total) still left the stability gate pass rate at 0/500. The root cause: score gaps between candidates were 7-10 orders of magnitude smaller than the noise floor, making most tasks fundamentally unrankable.
I think this is a great reminder that before using an LLM as a measurement instrument, you need to measure whether the instrument itself is stable.
#
LLMEval# #
Reproducibility#