Research
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
Overview Research area: Data management and information extraction (cs.DB) — specifically document-to-table (Doc2Table) extraction using large language models, plus benchmark construction via syntheti

- arXiv
- 2602.13812
- Published
- 2026-02-14
- Authors
- Yuxiang Guo, Zhuoran Du, Nan Tang, Kezheng Tang, Congcong Ge, Yunjun Gao
AI summary
Overview
- Research area: Data management and information extraction (cs.DB) — specifically document-to-table (Doc2Table) extraction using large language models, plus benchmark construction via synthetic data generation.
- Technical level: Intermediate. The paper is readable without deep database theory, but it assumes familiarity with LLM evaluation, schema-guided extraction, and metrics such as precision, recall, and F1.
- One-sentence scope: The paper introduces DTBench, a synthetic, capability-aware benchmark of 120 cases and 8,811 cell-level instances, built by reversing the usual extraction task (generating documents from known ground-truth tables) with a multi-agent pipeline, and uses it to evaluate eight mainstream LLMs across a two-level taxonomy of five capabilities and thirteen sub-capabilities.
What This Paper Is About
Organizations hold large amounts of unstructured text (clinical notes, contracts, financial filings), but reliable quantitative analysis needs structured tables that can be queried with SQL. Large language models are increasingly used to extract tables from documents under a target schema, yet it is unclear how well they handle cases where values must be transformed, reasoned about, filtered from distractors, left missing, or reconciled across conflicts. Existing benchmarks do not distinguish or cover these different capabilities, and building such a benchmark by hand is costly and hard to scale, so the authors generate documents backwards from curated ground-truth tables to create a benchmark that evaluates each capability systematically.
Key Contributions
- Doc2Table capability taxonomy: A two-level taxonomy of the capabilities required for document-to-table extraction, comprising 5 top-level categories — Transformative Alignment (TA), Reasoning & Inference (RI), Distractor Robustness (DR), Evidence Faithfulness (EF), and Conflict Resolution (CR) — and 13 sub-capabilities.
- Table2Doc synthesis workflow: A multi-agent, reverse-generation pipeline that treats capability as a latent variable so the synthesized documents cover the predefined taxonomy, and that adds multi-stage verification (deterministic code checkers plus checklist-based LLM verifiers) to enforce document quality.
- The DTBench benchmark: The first capability-aware benchmark for Doc2Table extraction, consisting of 120 synthesized cases with 8,811 cell-level evaluation instances, of which 4,518 (over 50%) carry capability labels, spanning documents of 498 to 182,297 tokens and ground-truth tables of 3–37 rows and 2–17 columns.
- Extensive experiments: Evaluations of eight LLMs on DTBench, an ablation study of the synthesis pipeline's agents, and fine-grained analysis of model strengths and weaknesses per capability and sub-capability, including per-document time and total inference cost.
Main Findings
- Benchmark coverage versus prior work: Table 1 compares existing benchmarks — Rotowire and E2E cover only Transformative Alignment, LiveSum covers Reasoning & Inference and Distractor Robustness, InstructIE covers Reasoning & Inference, Distractor Robustness, and Evidence Faithfulness, and StructText covers Transformative Alignment and Distractor Robustness. DTBench is the only listed benchmark marked as covering all five capability dimensions.
- Ablation shows each agent matters, differently: Removing the Annotator collapses Capability-awareness from 58.90% to 1.53%; removing the Verifier drops Completeness to 72.43% and Exclusiveness to 88.73% (from 100% for both in the full pipeline); removing the Planner lowers Capability-awareness to 39.17% and reduces linguistic scores (Textual Coherence 3.45 versus 4.10). The authors conclude that the Annotator controls dataset difficulty, the Refiner and Planner improve readability and coherence, and the Verifiers are essential for Completeness and Exclusiveness.
- Leading models still separate direct from indirect extraction: Overall F1 on DTBench ranges from 33.65 (Qwen3-4B) to 89.34 (Gemini-3-flash). The decomposed recall metrics reveal a consistent gap: for example, Gemini-3-flash reaches R_dir 95.93 versus R_ind 80.90 (a gap of 15.67%), GPT-5 reaches 92.48 versus 75.12 (18.77%), and Qwen3-4B reaches 35.47 versus 19.33 (45.50%).
- Sub-capability performance is highly uneven: In the fine-grained results, Rule-based Resolution includes scores of 100.00, while Missing Value Faithfulness includes scores as low as 0.53 and Multi-hop Reasoning includes scores as low as 7.73. The authors state that persistent challenges remain in reasoning, faithfulness, and conflict resolution.
- Cost and latency vary widely across models: Reported per-document times range from 10.68 seconds (Gemini-3-flash) to 84.32 seconds (GPT-5), and total costs range from 0.31 USD (Llama3.1-8B) to 23.92 USD (GPT-5), with Gemini-3-flash at 8.18 USD and GPT-5-mini at 4.65 USD.
- Model bias was tested, not just argued: To show that the evaluation and findings are not an artifact of the synthesizer, the authors rebuilt the complete dataset using GPT-4o-mini as an alternative synthesizer instead of the Grok-4-fast backbone used in the main pipeline, with results reported in the appendix.
- Table foundation models are a poor fit: The authors note that recent TableLLM-style foundation models are optimized for structured-table inputs, whereas DTBench provides unstructured documents, and they report an additional TableLLM experiment to verify this mismatch.
Methodology in Plain English
Rather than paying annotators to read documents and build ground-truth tables, the authors run the task in reverse. They start with 120 existing tables curated from Kaggle, Wikipedia, and public data-fusion datasets, each with an identifier column, and manually augment each table's schema with attribute descriptions, data types, and constraints. To test faithfulness specifically, they deliberately delete some cell values so those cells should be extracted as missing.
A five-stage multi-agent pipeline then writes a document that would produce exactly that table. In Stage 1, an LLM Annotator labels each cell with the capability needed to extract it (or marks it "empty" if directly extractable), and a code-based Checker re-prompts for any unannotated cells, defaulting to "empty" after three rounds. In Stage 2, a Refiner picks a specific sub-capability and inversely generates textual evidence for each cell value; a Verifier checks value correctness, label alignment, and schema leakage — whether the text accidentally enables extra extractable facts — and returns feedback for regeneration until it passes or a retry limit is hit. Stage 3 has a Planner choose a plausible document type and a section blueprint, checked by code for complete evidence coverage. Stage 4 has a Writer draft section by section, with an LLM Verifier checking faithful grounding and schema leakage and looping back for revisions. Stage 5 assembles the verified sections into the final document.
The authors use Grok-4-fast as the backbone for all agents. They argue the task is too hard for models to win by mimicking the synthesizer's style, and that the constrained, plan-driven generation limits the synthesizer's stylistic footprint. For evaluation, they align predicted tables to ground truth by treating row matching as weighted bipartite matching on key-attribute similarity, then compare cells after standard normalization and count exact matches. They report precision, recall, F1, type-specific recall for directly extractable (R_dir) and capability-demanding (R_ind) cells, time, and cost; a separate normalized sub-capability score is used for the fine-grained analysis.
Why This Matters
- Research impact: The paper reframes Doc2Table extraction as a set of distinct capabilities rather than one monolithic task, and provides a scalable construction recipe (reverse generation from ground-truth tables) that other benchmark builders can reuse. It also gives the community a testbed with explicit cell-level labels, so failures can be attributed to specific abilities rather than reported as a single accuracy number.
- Real-world applications:
- Healthcare analytics — turning clinical notes, electronic health reports, and administrative records into tables that support deterministic SQL queries such as complication rates for patients on a given drug.
- Financial and legal document processing — extracting structured fields from financial filings, contracts, and exhibits where unit scaling, date formats, and original-versus-amended dates matter.
- Auditable business intelligence — producing tables that let auditors trace which evidence supports each computed number, unlike opaque RAG answers.
- Multi-source data fusion — reconciling conflicting values for the same attribute reported by different sources, as in stock reporting, where source-aware resolution is required.
- Industry relevance: The reported cost and latency numbers show that even the strongest models here are inexpensive per document in absolute terms (0.31–23.92 USD total over the benchmark), but that the accuracy gap between direct and indirect extraction — up to 45.50% — is the practical obstacle to deploying LLM-based extraction in reliability-critical pipelines. The finding that removing verification hurts Completeness and Exclusiveness also matters for teams designing their own synthetic data.
Future Directions
- Closing the indirect-extraction gap: Every model evaluated shows lower recall on capability-demanding cells than on directly extractable cells, with gaps from 15.67% to 45.50%; improving reasoning, faithfulness, and conflict resolution is the central open problem.
- Improving missing-value behavior: Missing Value Faithfulness produced the lowest reported sub-capability scores, indicating that models often invent plausible values instead of emitting NULL, a failure with direct downstream consequences.
- Scaling and diversifying synthesis: The pipeline depends on Grok-4-fast as the backbone and on 120 curated tables; testing other synthesizers (as done with GPT-4o-mini) and broader table sources would clarify how much benchmark difficulty is synthesizer-dependent.
- Adapting table-specialized models: Table foundation models currently perform poorly because they expect structured-table input; whether they can be adapted for unstructured documents remains an open question the authors flag.
Target Audience
Researchers and practitioners working on information extraction, LLM evaluation, data integration and fusion, and document understanding — particularly those building or auditing pipelines that convert unstructured documents into queryable tables. It is also useful for benchmark designers interested in synthetic, capability-aware dataset construction, and for teams deciding which LLM to deploy for extraction given trade-offs among accuracy, latency (10.68–84.32 seconds per document), and cost (0.31–23.92 USD).
Authors’ abstract
Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and verifiable SQL-based data analytics. Although large language models (LLMs) have shown promise in flexible information extraction, their ability to produce precisely structured tables remains insufficiently understood, particularly for indirect extraction that requires complex capabilities such as reasoning and conflict resolution. Existing benchmarks neither explicitly distinguish nor comprehensively cover the diverse capabilities required in Doc2Table extraction. We argue that a capability-aware benchmark is essential for systematic evaluation. However, constructing such benchmarks using human-annotated document-table pairs is costly, difficult to scale, and limited in capability coverage. To address this, we adopt a reverse Table2Doc paradigm and design a multi-agent synthesis workflow to generate documents from ground-truth tables. Based on this approach, we present DTBench, a synthetic benchmark that adopts a proposed two-level taxonomy of Doc2Table capabilities, covering 5 major categories and 13 subcategories. We evaluate several mainstream LLMs on DTBench, and demonstrate substantial performance gaps across models, as well as persistent challenges in reasoning, faithfulness, and conflict resolution. DTBench provides a comprehensive testbed for data generation and evaluation, facilitating future research on Doc2Table extraction. The benchmark is publicly available at https://github.com/ZJU-DAILY/DTBench.