Introducing ExtractBench on Kaggle Benchmarks with
@llama_index.
When AI agents rely on schema-guided extraction before human review, one truncated schedule or invented value becomes a wrong payment or decision.
ExtractBench evaluates models in workflows based on real-world documents across industries such as supply chain, healthcare, and finance, by measuring whether the system returns:
➣ Missing fields as null instead of inventing a value
➣ Source evidence for each value
➣ Every record of each repeated structure
GPT-5.6 Sol currently leads at 91%.