๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
279 ํŒ”๋กœ์ž‰ ์ค‘    414 ํŒฌ
๐Ÿงญ TL;DR: Instead of estimating an LLM's confidence from the current inference alone, this method calibrates it against how often similar past attempts actually succeeded. It beats 10-sample self-consistency in accuracy while costing far less. Title: Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents URL: Points ๐Ÿ—‚ Stores past task, reasoning trace, stated confidence, actual outcome, and a lesson learned in an experience bank ๐Ÿ” Recall stage retrieves k=50 similar episodes and computes their real hit rate statistically ๐Ÿ’ญ Reflect stage has the model restate confidence in words after seeing that track record ๐Ÿ† Matches or beats SC@10 on 23 of 24 model-dataset combinations by AUROC ๐Ÿค– Biggest gains on agent tasks where failures are silent โ€” even beats a trained verifier on AppWorld ๐Ÿ“‰ 2-20x lower calibration error across domains, at roughly 1/10th the compute cost A great example of the shift from "judge the current attempt" to "judge from accumulated track record." #LLMEval# #AIAgents#
๋” ๋ณด๊ธฐ