Skip to content
AI.info

Research

SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats

Overview Research area: Natural Language Processing / retrieval-augmented generation (RAG) for question answering over spreadsheets and tabular data, with an emphasis on financial documents. Technical

SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
arXiv
2512.04292
Published
2025-12-03
Authors
Chinmay Gondhalekar, Urjitkumar Patel, Fang-Chun Yeh

AI summary

Overview

Research area: Natural Language Processing / retrieval-augmented generation (RAG) for question answering over spreadsheets and tabular data, with an emphasis on financial documents.

Technical level: Intermediate. The paper assumes familiarity with RAG, vector embeddings, and SQL, but the routing logic and evaluation protocol are described in accessible terms.

Scope: The paper proposes SQuARE, a hybrid retrieval framework that inspects a spreadsheet's structural complexity and routes each natural-language question either to structure-preserving chunk retrieval or to constrained SQL over an automatically built relational view, with an agentic fallback that switches or merges paths when confidence is low.

What This Paper Is About

Real-world spreadsheets — especially corporate financial statements — are hard to query with language models because multi-row headers, merged cells, and unit annotations get broken apart by naive text chunking, while pure SQL approaches fail on files that lack a consistent, tidy schema. SQuARE's goal is to stop treating every table the same: it measures how structurally messy a sheet is and sends each question down the retrieval path that is safest for that sheet, returning the exact cells or rows used as evidence.

Key Contributions

  1. Complexity metric for routing. A score based on header depth (H) and the count of merged or split cells in the header region (M), computed as X = αH + βM and applied with a sheet-normalized threshold, distinguishes "Multi-Header" sheets from "Flat" sheets so that only the necessary backend is built.
  2. Structure-preserving segmentation with semantic indexing. For complex sheets, the algorithm segments at header boundaries and attaches metadata capturing the ordered header path, time labels, and unit strings; it embeds short two-sentence block descriptions rather than raw cells.
  3. Schema-aware SQL generation with guardrails. For flat tables, SQuARE infers a cleaned schema with column names and types, persists detected units, and generates constrained SQL restricted to SELECT–FROM–WHERE–GROUP BY–ORDER BY–LIMIT over a whitelisted schema, with a strict row cap and no DDL/DML.
  4. Agentic routing with confidence-aware fallback. A lightweight agent picks between chunk and SQL based on sheet label and query cues; on low confidence it switches modes or merges contexts, summarizes if the token budget is exceeded, and abstains if the merged context still fails the quality gate.

Main Findings

  • Multi-header corporate balance sheets: SQuARE(Gemma) reached 91.3% accuracy, versus SQuARE(Llama) at 80.7% and ChatGPT-4o at 28.7% on 400 QA pairs drawn from ten issuers (Microsoft, Meta, Alphabet, Netflix, Tesla, Adobe, SAP, Nvidia, Amazon, Dell Technologies).
  • Merged World Bank workbook: SQuARE(Gemma) reached 86.0%, ahead of SQuARE(Llama) at 74.0% and ChatGPT-4o at 54.0% on 50 QA pairs.
  • Flat single-header tables: Across 450 QA pairs (five datasets, 30 Easy / 30 Medium / 30 Hard each), SQuARE(Gemma) achieved 93.3% overall and 87% on the Hard tier, while ChatGPT-4o trailed on Hard at 81%.
  • Retrieval recall: On the 10 multi-header balance sheets (400 Q), FAISS R@3 was 0.88; on the merged World Bank workbook (50 Q) it was 0.86. On flat tables, SQL R@1 was 0.95 (Easy), 0.92 (Medium), and 0.90 (Hard), with FAISS R@3 at 0.87 / 0.85 / 0.81 and merged R@3 at 0.89 / 0.88 / 0.84.
  • Ablation results: Removing fallback cost 4–6 points across settings (Multi-Header 91.3% → 89.0%; Merged WB 86.0% → 80.0%; Flat 93.3% → 90.0%). Chunk-only degraded flat tables sharply to 75.5%, and SQL-only on flat tables reached 90.0%.
  • Routing behavior: Across 400 multi-header queries, SQuARE chose chunking 100% of the time (fallback under 8%). For flat tables, the agent selected chunking 35% and SQL 65%, with an overall fallback rate of roughly 10%.
  • Cost and latency: On a reference setup of quantized models on a T4 GPU with 15 GB VRAM, end-to-end latency was close to ChatGPT-4o, slightly slower on complex financial spreadsheets and nearly comparable on flat files.
  • Error analysis: ChatGPT-4o failures were attributed to confusion among multiple relevant numerical entries, unit mismatches, and an inability to reference specific table rows and columns.
  • Hyperparameters: α = 0.6, β = 0.4, at most k = 3 forwarded chunks, merge density threshold ρ ∈ [0.10, 0.15], temperature 0, and no dataset-specific tuning.

Methodology in Plain English

The researchers built a system that first looks at a spreadsheet and decides how complicated its layout is. Two things are measured: how many rows of headers are stacked on top of each other, and how many merged cells sit in the header area. These are combined into a single score, and a sheet-normalized cutoff sorts sheets into "Multi-Header" and "Flat."

For every sheet, a semantic index is built. On Multi-Header sheets, the file is cut into blocks at header boundaries rather than at fixed character windows, and each block carries metadata about its header path, the years it covers, and any unit line such as "USD (millions)." Each block gets a short two-sentence description that is embedded into a vector index.

On Flat sheets, the system additionally constructs a small relational view, infers column names and types, and can generate restricted SQL queries. When a question arrives, a lightweight agent picks a path based on the sheet label and question wording — filter-and-aggregate questions lean toward SQL on flat sheets, layout-referencing questions lean toward chunk retrieval. The chosen context must pass a quality check; if it fails, the system tries the other path on flat sheets, merges the two sets of evidence, summarizes if the merged context exceeds the token budget, and otherwise abstains rather than guessing. Answers come with the specific cells or rows used.

Evaluation used exact-match accuracy for numeric and categorical answers, plus retrieval recall reported separately for the chunk and SQL paths, so retrieval quality is isolated from answer generation. All outputs were manually verified, and the ChatGPT-4o baseline was run through the public web interface with no tools, browsing, or external retrieval.

Why This Matters

Impact on research. The paper argues against one-size-fits-all retrieval for tables and shows that a structural complexity signal can drive per-query routing. It connects spreadsheet QA to related lines of work such as TabRAG, TableRAG, fine-tuned tabular embedding models, NL2SQL, and the emerging tabular foundation model track (TaBERT, TAPAS, TABBIE, TabPFN/TabPFNv2, TabICL, TabDPT), positioning SQuARE as a bridge between them.

Real-world applications:

  • Financial statement analysis over corporate balance sheets with nested headers, unit lines, and fiscal-versus-calendar year labels.
  • Cross-indicator development analysis over merged World Bank workbooks such as Gender Statistics combined with World Development Indicators.
  • Public-sector and macroeconomic reporting on flat datasets like Public Sector Debt, Global Economic Prospects, Energy Consumption, and Education Attainment.
  • Any audit or compliance workflow where an answer must be traceable to the exact spreadsheet cell or row that produced it.

Industry relevance. The system is modular and runs on modest hardware — quantized models on a T4 GPU with 15 GB VRAM — and its embeddings, vector stores, and databases are swappable without changing control flow. Caching and lazy indexing (building the relational view only when needed) keep runtime predictable, which matters for deployment cost. The authors are affiliated with Ratings Data Science at S&P Global and the paper was accepted to the IEEE International Workshop on Large Language Models in Finance (December 8–11, Macau, China, 2025), underscoring direct financial-industry interest.

Future Directions

  • A learned router. The current agentic router is prompt-based; a lightweight learned router with uncertainty estimates could reduce fallbacks and sharpen confidence checks, possibly under a cost-aware objective balancing accuracy, context length, and latency.
  • Layout and format robustness. The experiments focused on well-structured spreadsheets. Handling OCR'd tables and document-style layouts would require integrating table detectors and layout parsers ahead of the same retrieval backbone, potentially following MultiFinRAG's multimodal extraction and tiered fusion while keeping SQuARE's SQL/structure routing.
  • SQL robustness and join scope. SQL generation can be fragile under schema drift, so the authors plan schema alignment, column–header aliasing, and safe join discovery, plus multi-sheet and cross-workbook queries with explicit, verifiable joins.
  • Evaluation coverage and model recency. The authors note they used manual grading and a public, tool-free ChatGPT-4o comparison, and call for a broader baseline set including tool-augmented assistants, perturbation robustness tests, and public release of de-identified QA pairs where licensing permits. They also flag that GPT-5 and GPT-oss became available during finalization and were only smoke-tested on a small slice.

Target Audience

This paper is most useful to NLP and RAG engineers building question-answering systems over structured data, data scientists and analysts in financial services who work with real-world spreadsheets, and researchers studying table understanding, NL2SQL, or tabular foundation models. Practitioners evaluating retrieval architectures for enterprise document and spreadsheet workflows will find the routing logic, ablation results, and invocation budget analysis directly actionable.

Authors’ abstract

Accurate question answering over real spreadsheets remains difficult due to multirow headers, merged cells, and unit annotations that disrupt naive chunking, while rigid SQL views fail on files lacking consistent schemas. We present SQuARE, a hybrid retrieval framework with sheet-level, complexity-aware routing. It computes a continuous score based on header depth and merge density, then routes queries either through structure-preserving chunk retrieval or SQL over an automatically constructed relational representation. A lightweight agent supervises retrieval, refinement, or combination of results across both paths when confidence is low. This design maintains header hierarchies, time labels, and units, ensuring that returned values are faithful to the original cells and straightforward to verify. Evaluated on multi-header corporate balance sheets, a heavily merged World Bank workbook, and diverse public datasets, SQuARE consistently surpasses single-strategy baselines and ChatGPT-4o on both retrieval precision and end-to-end answer accuracy while keeping latency predictable. By decoupling retrieval from model choice, the system is compatible with emerging tabular foundation models and offers a practical bridge toward a more robust table understanding.

Read the original paper