注册并分享邀请链接,可获得视频播放与邀请奖励。

Jerry Liu
@jerryjliu0
Parsing the world's hardest PDFs @llama_index. cofounder/CEO Careers: Enterprise:
加入 September 2011
1.6K 正在关注    84.6K 粉丝
predicting model uncertainty is a hard problem. perfectly "calibrated" confidence scores are exact probability values on whether the output is correct. this is extremely important for agentic decision making, including document extraction. i made a sick video below showing how confidence scores can be used to choose decision thresholds and vary precision / recall. if you set a really high threshold, then you automate less, but more of the automated extraction is correct. if you set a low threshold, then you automate more, but there's more errors in the extraction. check out our blog!
显示更多
confidence scores only matter if they help you decide what to automate. for document extraction, that usually means knowing how much work you can safely accept at a given precision target. in our latest post, we look at confidence scoring through that lens, including: ✅️ confidence cutoffs ✅️ precision vs. recall ✅️ score coverage ✅️ score granularity ✅️ human review volume using ExtractBench, we compare how different extraction systems perform after confidence filtering. at a 97% precision target, LlamaParse Agentic Plus reached 66.48% recall on expected fields after filtering. the useful part of a confidence score isn’t the number itself. it’s whether you can use it to control automation and review in production. 👉️ read the full post:
显示更多
0
18
79
11
转发到社区