注册并分享邀请链接,可获得视频播放与邀请奖励。

Jerry Liu
@jerryjliu0
Parsing the world's hardest PDFs @llama_index. cofounder/CEO Careers: Enterprise:
加入 September 2011
1.6K 正在关注    84.6K 粉丝
This is a fantastic blog post from our research team on the importance of calibrated confidence scores. Here it's in the context of document extraction, it's also extremely important for more general agentic decision making (eg with Jev) When you set a confidence threshold, you can choose to automatically accept values above that threshold and do HITL review of values below the threshold. The higher the confidence threshold, the higher precision you're able to guarantee (e.g. confidence of 0.7 could mean 95% precision, confidence of 0.9 could mean 98% precision), but of course the more human review you'd have to do on false negatives. We've put in a lot of work to make sure our confidence scores are well calibrated and represents real uncertainty over complex documents in production. Come check out our blog: LlamaParse:
显示更多
confidence scores only matter if they help you decide what to automate. for document extraction, that usually means knowing how much work you can safely accept at a given precision target. in our latest post, we look at confidence scoring through that lens, including: ✅️ confidence cutoffs ✅️ precision vs. recall ✅️ score coverage ✅️ score granularity ✅️ human review volume using ExtractBench, we compare how different extraction systems perform after confidence filtering. at a 97% precision target, LlamaParse Agentic Plus reached 66.48% recall on expected fields after filtering. the useful part of a confidence score isn’t the number itself. it’s whether you can use it to control automation and review in production. 👉️ read the full post:
显示更多
0
20
51
4
转发到社区