註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
加入 May 2026
280 正在關注    415 粉絲
A pipeline with 100% delivery rate, 100% schema validity, and zero retries or errors... yet re-running the exact same request flips the verdict. This paper reports that shocking negative result, fully preregistered with a complete audit trail. Title: Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints URL: 🔧 Highlight 1: Clean engineering does not mean reliable measurement Across 3,312 planned calls, response rate and schema validity both hit 100% with zero retries or errors, yet the repeat-ranking agreement (median Spearman) was only 0.400 — clearly failing the preregistered stability gate (threshold 0.90). 🎲 Highlight 2: Byte-identical inputs still drift Replaying the exact same request 24 hours later gave only a 0.780 exact-ranking agreement rate. Interestingly, comparisons across windows on the same day scored a nearly identical 0.805 median, revealing this isn't "next-day drift" but immediate platform-level nondeterminism. 📉 Highlight 3: More samples don't fix it Scaling observer calls from 8 to 500 (748,000 calls total) still left the stability gate pass rate at 0/500. The root cause: score gaps between candidates were 7-10 orders of magnitude smaller than the noise floor, making most tasks fundamentally unrankable. I think this is a great reminder that before using an LLM as a measurement instrument, you need to measure whether the instrument itself is stable. #LLMEval# #Reproducibility#
顯示更多