Inside ExtractBench: How to Evaluate Document Extraction for AI Agents
26
4:00 PM - 4:45 PM
Most document extraction systems look great on short, clean PDFs, and quietly fall apart in production. When our applied research team benchmarked 14 systems (frontier VLMs, coding agents, and specialized extraction APIs) against 370 enterprise documents, the results were stark: on files past 50 pages, commercial VLMs collapse below 35% recall, silently dropping most table rows while precision stays deceptively high.
Join Simon Suo, CTO and co-founder of LlamaIndex, for a deep dive into ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents: 4,869 pages across 67 document types, evaluated with zero LLM judges for fully deterministic, reproducible results.
In this session, Simon will cover:
- How the benchmark was built — why we chose value accuracy, long-record completeness, spatial grounding, and per-page cost as the metrics that matter for production, and how we curated the dataset behind them
- Where extraction methods fail — the failure modes that short-document benchmarks hide, from silent list truncation to missing spatial citations
- How to hill climb — practical strategies for evaluating your own extraction pipeline against state-of-the-art benchmarks like ExtractBench and systematically improving extraction quality
Whether you're building agents that depend on reliable document data or evaluating extraction vendors, you'll leave knowing exactly how to measure, and improve, the extraction quality of your systems.
Speaker
Simon Suo
CTO, LlamaIndex
26
4:00 PM - 4:45 PM