Skip to content
AI.info

Research

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation

Overview Research area: Natural Language Processing / evaluation of Large Language Models on table question answering, applied to telecommunications technical standards (3GPP specifications). Technica

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
arXiv
2601.04202
Published
2025-12-05
Authors
Anas Ezzakri, Nicola Piovesan, Mohamed Sana, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang

AI summary

Overview

Research area: Natural Language Processing / evaluation of Large Language Models on table question answering, applied to telecommunications technical standards (3GPP specifications).

Technical level: Advanced. The paper assumes familiarity with LLM prompting and evaluation metrics (pass@1, cons@16), retrieval-augmented generation, supervised fine-tuning, and reinforcement learning with verifiable reward.

Scope: The paper builds and releases TeleTables, a benchmark of 2,220 tables from 13 3GPP specifications and 500 human-verified multiple-choice questions, then uses it to evaluate 20 open-weight LLMs in two settings: without the table (closed-book) and with the table provided as context.

What This Paper Is About

3GPP standards encode much of their technical content in complex tables with multi-level headers, merged cells, footnotes, and nested formulas, and prior telecom benchmarks such as TeleQnA found that standards-based questions are the hardest category for LLMs. The paper asks whether this difficulty comes from missing domain knowledge or from an inability to interpret the tables themselves. It answers this by releasing a multi-format table corpus and a human-verified MCQ benchmark, then measuring models both with and without the relevant table in context.

Key Contributions

  1. Table collection. A public corpus of 2,220 tables from 13 3GPP specifications (Release 18–19, spanning the TS 23.xx, 36.xx, and 38.xx series) released in four synchronized formats: HTML, JSON, Markdown, and PNG image. The authors describe it as the first multi-format collection of tables from a technical standards domain.
  2. Benchmark. TeleTables, consisting of 500 human-verified multiple-choice questions paired with their source tables, of which 250 are classified as Difficult (requiring multi-step or aggregated reasoning). Questions are annotated by reasoning skill (Q-Lookup, Q-Count, Q-Arithmetic, Q-Compare, Q-Average, Q-MultiCriteria) and evidence scope (E-Cell, E-Cross, E-Full).
  3. Pipeline. A two-stage, LLM-based MCQ generation framework combining multimodal generation, multi-trial validation, a difficulty/simplicity filter, and human review by 8 expert annotators with 3GPP experience. The framework is presented as reusable for benchmark construction from any technical document corpus.
  4. Findings. An evaluation of 20 open-weight LLMs across non-reasoning, multimodal, reasoning, and table-specialized architectures, identifying two distinct bottlenecks (domain knowledge closed-book, table interpretation with-table) and showing that non-telecom table specialization does not consistently help.

Main Findings

  • Closed-book performance is low. With no table and no internet access, models average roughly 35% cons@16, from Llama3.1-8B-Instruct at 17% to GPT-OSS-120B at 40.6%. No general-purpose model exceeds 40.6% cons@16.
  • Domain fine-tuning beats scale in the closed-book setting. A Qwen2.5-7B-Instruct model fine-tuned on 3GPP material rises from 26.15% to 50.48% pass@1 and from 26.00% to 58.00% cons@16, outperforming GPT-OSS-120B (40.60% cons@16) despite being approximately 17x smaller.
  • With the table in context, accuracy jumps above 90% for the best models. Qwen3-32B reaches 91.18% pass@1 and 92.60% cons@16 overall, and 83.60% pass@1 / 86.40% cons@16 on the Difficult subset.
  • HTML is the strongest single representation for most models, attributed to explicit markup of hierarchical cell relationships. Llama models are an exception, performing marginally better with Markdown. Multi-format concatenation gives modest gains for small models (+5.2pp for Llama3.1-8B with JSON+MD) and negligible improvement for larger ones (+0.6pp for Qwen3-32B).
  • Images are the weakest format. Qwen2.5-VL-72B drops from 85.6% with HTML to 77.2% with images; Gemma-27B-it drops to 53.8%. Hybrid image+text inputs recover most of the loss but at 12,831 tokens.
  • Markdown is the most efficient format, using 42% fewer tokens than HTML (707 vs 1,224 average tokens per table) at competitive accuracy.
  • Reasoning models dominate. All other reasoning models stay above 85% cons@16 overall, and even the small Qwen3-4B surpasses most larger non-reasoning models.
  • Multimodal models outperform their text-only counterparts even without images. Qwen2.5-VL-72B reaches 84.46% pass@1 / 87.40% cons@16 with HTML+JSON, versus 76.66% / 79.40% for text-only Qwen2.5-72B.
  • Non-reasoning models scale with size but struggle on hard questions. Models below 10B parameters fail to exceed 75% cons@16; Qwen2.5-32B reaches 85.60% cons@16 overall but drops to 72.00% on the Difficult subset.
  • Table specialization on non-telecom data provides no consistent benefit. TableGPT2-7B scores below its Qwen2.5-7B base (46.40% vs. 49.40% cons@16); TableGPT-R1 and STaR-8B, both built on Qwen3-8B, score below their unmodified base (87.20% and 88.20% vs. 89.60%). Table-R1-Zero-7B, trained with RLVR on the same 7B base as TableGPT2-7B, instead gains +23.6pp (73.00% cons@16).
  • Difficulty is structural and systematic. For Qwen3-32B, accuracy spans 32.2pp across reasoning skills, from 94.7% on direct lookup to 62.5% on multi-criteria questions; evidence scope falls from 93.2% (single cell) to 86.2% (full-table scan); and table structure falls from 100.0% on small tables (≤5×4) to 88.0% on large tables (>20 rows / 8 columns).
  • Complexity degrades performance for all model types. Using HTML token count as a proxy, Qwen3-32B achieves 100% pass@1 on the simplest tables (fewer than 500 tokens) and drops to about 80% on the most complex; Qwen2.5-7B-Instruct stays below 80% even on simple tables and degrades to 27% on the most complex.
  • Two bottlenecks, not one. Closed-book, domain knowledge is the primary constraint; with the table provided, interpretation becomes the primary constraint. The domain fine-tuned model reaches only 65.00% cons@16 with the table, versus 92.60% for the best general-purpose reasoning model.
  • Corpus complexity. Of the 2,220 tables, 28.4% have multi-level headers, 13.1% contain grouped rows with merged cells, and 38.3% include footnotes. Content is heterogeneous: 76.2% numeric/parametric, 43.0% prose descriptions, 34.2% categorical labels/codes, 30.8% mathematical expressions, and 2.7% symbol or bit sequences.

Methodology in Plain English

The authors first harvested tables from 13 3GPP specification documents. A four-stage pipeline localized tables by their standard title convention, extracted them with Pandoc into HTML and derived JSON and Markdown from that, rendered high-resolution PNG images preserving merged cells and hierarchical headers, attached metadata such as identifier, caption, and surrounding paragraphs, and used Qwen2.5-VL-72B-Instruct to transcribe non-textual elements and verify fidelity against the image.

Questions were then generated in two stages. Stage one used a generator agent (Qwen2.5-VL-72B-Instruct) given all four table representations to write basic MCQs, with answer choices randomly shuffled to reduce positional bias; a second agent answered each question blind across four trials, and a question was kept only if the correct option was chosen in at least three trials. Stage two used a reasoning model (Qwen3-32B) to synthesize five randomly selected basic questions per table into harder, multi-step questions, validated by Qwen2.5-72B-Instruct over four trials, with 0 successful trials discarding the item, 1–2 flagging it for human review, and 3–4 retaining it. A simplicity filter using Qwen2.5-7B-Instruct discarded questions a small model answered consistently. The funnel ran from 3,000 generated MCQs to 1,412 passing validator filtering to 817 passing human review to 500 sampled for release, with review by 8 expert annotators who required factual correctness, expert plausibility, and exactly one defensible option.

Evaluation used two settings. In the closed-book setting, the prompt referred to the table by identifier and document title but provided no table and no internet access; each model produced N=16 responses per question at temperature 0.6, top-p 0.90, and maximum generation length 16,384 tokens, scored with pass@1 and cons@16 (majority vote across 16 answers). In the with-table setting, the relevant table was supplied in one or more formats, simulating perfect retrieval as an upper bound on RAG performance. Additional experiments fine-tuned Qwen2.5-7B-Instruct on 3GPP tables and documents for table QA, table completion, structure extraction, and multi-step reasoning, excluding all TeleTables benchmark questions and answers, and stratified results by reasoning skill, evidence scope, table size, row organization, and HTML token count.

Why This Matters

Impact on research. The paper separates two failure modes that prior telecom benchmarks conflated: weak domain knowledge versus weak table interpretation. It shows that the second persists even when the correct table is in context, which means retrieval-augmented systems for standards documents face an unreduced bottleneck. It also provides evidence that table specialization on general-domain corpora does not transfer to technical standards, and that training objective and data distribution may matter more than specialization itself.

Real-world applications.

  • Automated configuration of network parameters derived from 3GPP tables.
  • Parameter look-up and cross-referencing by network engineers.
  • Specification-driven troubleshooting, where a fault must be traced back to a conditional table entry.
  • Retrieval-augmented assistants that answer questions directly from standards documents.

Industry relevance. Telecom operators and equipment vendors rely on 3GPP specifications that contain hundreds of tables per document, often with nested headers, merged cells, cross-table references, and condition-dependent entries. The paper's finding that a 7B model fine-tuned on 3GPP data beats far larger general-purpose models closed-book, but is still far behind a strong reasoning model when the table is available, gives a concrete build-versus-buy signal for telecom LLM deployment: domain adaptation addresses knowledge gaps, while reasoning capability addresses interpretation.

Future Directions

  • Extending coverage beyond the 13 3GPP specifications from Releases 18 and 19 to other releases, other standardization bodies such as IETF and IEEE 802, and other technical domains.
  • Building cross-lingual versions of the benchmark; all source documents and MCQs are currently in English.
  • Increasing benchmark size beyond 500 MCQs by sampling from a broader set of specifications to capture more of the diversity of 3GPP tables.
  • Studying how pretraining gaps interact with retrieval quality in end-to-end RAG pipelines, since the current with-table experiment assumes perfect retrieval.
  • Closing the interpretation gap: determining what training or architecture changes would let models handle multi-criteria reasoning, full-table evidence scopes, and large complex tables at the level reasoning models reach on simple lookups.

Target Audience

Researchers working on table question answering, LLM benchmarking, and retrieval-augmented generation; telecom engineers and standards specialists interested in automating work over 3GPP specifications; and practitioners deciding whether to invest in domain fine-tuning, table-specialized models, or reasoning-capable general models for document-heavy technical workflows. Readers should be comfortable with LLM evaluation methodology and standard fine-tuning and reinforcement-learning terminology.

Authors’ abstract

Large Language Models (LLMs) are increasingly applied to telecom engineering tasks, yet perform poorly on 3GPP specifications. These standards encode much of their technical information in complex tables, but LLM knowledge and interpretation of such tables remain largely unexplored. We introduce TeleTables, a benchmark comprising 2,220 tables from 13 3GPP specifications in four formats and 500 human-verified MCQs spanning direct retrieval to multi-step reasoning. Evaluating 20 open-weight LLMs across non reasoning, multimodal, reasoning, and table specialized architectures reveals two distinct performance bottlenecks. In the closed-book setting, domain knowledge is the primary constraint, with no general-purpose model exceeding 41% accuracy. When the table is provided as context, the best models exceed 90%, but performance degrades systematically with reasoning depth, evidence scope, and structural complexity, with a 32.2pp spread across reasoning skills. Table specialization on non-telecom data provides no consistent benefit, while strong reasoning capabilities remain essential for reliable interpretation of complex technical tables.

Read the original paper