๐งญ TL;DR: Instead of estimating an LLM's confidence from the current inference alone, this method calibrates it against how often similar past attempts actually succeeded. It beats 10-sample self-consistency in accuracy while costing far less.
Title: Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
URL:
Points
๐ Stores past task, reasoning trace, stated confidence, actual outcome, and a lesson learned in an experience bank
๐ Recall stage retrieves k=50 similar episodes and computes their real hit rate statistically
๐ญ Reflect stage has the model restate confidence in words after seeing that track record
๐ Matches or beats SC
@10 on 23 of 24 model-dataset combinations by AUROC
๐ค Biggest gains on agent tasks where failures are silent โ even beats a trained verifier on AppWorld
๐ 2-20x lower calibration error across domains, at roughly 1/10th the compute cost
A great example of the shift from "judge the current attempt" to "judge from accumulated track record."
#
LLMEval# #
AIAgents#