🧭 TL;DR: Instead of estimating an LLM's confidence from the current inference alone, this method calibrates it against how often similar past attempts actually succeeded. It beats 10-sample self-consistency in accuracy while costing far less.
Title: Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
URL:
Points
🗂 Stores past task, reasoning trace, stated confidence, actual outcome, and a lesson learned in an experience bank
🔍 Recall stage retrieves k=50 similar episodes and computes their real hit rate statistically
💭 Reflect stage has the model restate confidence in words after seeing that track record
🏆 Matches or beats SC
@10 on 23 of 24 model-dataset combinations by AUROC
🤖 Biggest gains on agent tasks where failures are silent — even beats a trained verifier on AppWorld
📉 2-20x lower calibration error across domains, at roughly 1/10th the compute cost
A great example of the shift from "judge the current attempt" to "judge from accumulated track record."
#
LLMEval# #
AIAgents#