๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
279 ํŒ”๋กœ์ž‰ ์ค‘    414 ํŒฌ
A pipeline with 100% delivery rate, 100% schema validity, and zero retries or errors... yet re-running the exact same request flips the verdict. This paper reports that shocking negative result, fully preregistered with a complete audit trail. Title: Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints URL: ๐Ÿ”ง Highlight 1: Clean engineering does not mean reliable measurement Across 3,312 planned calls, response rate and schema validity both hit 100% with zero retries or errors, yet the repeat-ranking agreement (median Spearman) was only 0.400 โ€” clearly failing the preregistered stability gate (threshold 0.90). ๐ŸŽฒ Highlight 2: Byte-identical inputs still drift Replaying the exact same request 24 hours later gave only a 0.780 exact-ranking agreement rate. Interestingly, comparisons across windows on the same day scored a nearly identical 0.805 median, revealing this isn't "next-day drift" but immediate platform-level nondeterminism. ๐Ÿ“‰ Highlight 3: More samples don't fix it Scaling observer calls from 8 to 500 (748,000 calls total) still left the stability gate pass rate at 0/500. The root cause: score gaps between candidates were 7-10 orders of magnitude smaller than the noise floor, making most tasks fundamentally unrankable. I think this is a great reminder that before using an LLM as a measurement instrument, you need to measure whether the instrument itself is stable. #LLMEval# #Reproducibility#
๋” ๋ณด๊ธฐ