Skip to content
AI.info

Research

TelcoAI: Advancing 3GPP Technical Specification Search through Agentic Multi-Modal Retrieval-Augmented Generation

Overview Research area: Domain-specific (telecommunications) retrieval-augmented generation, combining agentic query planning with multi-modal document understanding for 3GPP standards documents. Tech

TelcoAI: Advancing 3GPP Technical Specification Search through Agentic Multi-Modal Retrieval-Augmented Generation
arXiv
2601.16984
Published
2025-11-17
Authors
Rahul Ghosh, Chun-Hao Liu, Gaurav Rele, Vidya Sagar Ravipati, Hazar Aouad

AI summary

Overview

Research area: Domain-specific (telecommunications) retrieval-augmented generation, combining agentic query planning with multi-modal document understanding for 3GPP standards documents.

Technical level: Advanced.

Scope: This paper presents TelcoAI, a multi-stage RAG system that ingests 3GPP technical specifications in .docx format and answers complex engineering queries by combining section-aware chunking, metadata-guided retrieval, agentic query decomposition, and text-plus-diagram fusion.

What This Paper Is About

3GPP technical specifications are long, deeply hierarchical, dense with equations, tables, diagrams, and cross-references, which makes them hard for engineers to search and interpret. Large language models alone struggle with these documents, and existing RAG systems for the domain typically use flat chunking, ignore the visual content, and rely on single-step retrieval over simple, isolated evaluation queries. The paper's goal is to build a retrieval and answering system that respects the document structure, handles diagrams and tables as well as text, and decomposes complex multi-document, version-sensitive questions into a planned sequence of retrievals.

Key Contributions

  1. Agentic multi-stage RAG architecture: A modular retrieval and generation pipeline integrating hierarchical reasoning, query planning, and answer synthesis.
  2. Multi-modal document understanding: A custom ingestion pipeline that fuses visual content (technical diagrams and tables) with text using metadata-aware alignment, by embedding LLM-generated natural-language descriptions and element markers into chunks.
  3. Section-aware chunking and structured retrieval: A recursive, section-based decomposition that preserves the document hierarchy, keeping semantically related content cohesive across retrieved contexts.
  4. Agentic query decomposition and fusion: A planner that splits complex queries into sub-queries and follow-up queries, plus answer fusion over content slices retrieved from multiple chunks.
  5. Comprehensive evaluation: Experiments on curated and synthetic datasets for the 3GPP domain, which the authors say they are working on releasing.

Main Findings

  • Overall performance: TelcoAI achieves 87% recall, 83% claim recall, and 92% faithfulness, described as a 16% improvement over state-of-the-art baselines.
  • Human-QA results: TelcoAI with Claude 3.5 Sonnet v2 reaches Recall 0.87 and Claim Recall 0.83, versus the best baseline (Chat3GPP with Claude 3.5 Sonnet v2) at 0.75 and 0.73 — a 16% and 13.7% improvement respectively.
  • Synthetic-QA results: TelcoAI reaches Recall 0.85 and Claim Recall 0.83, versus Chat3GPP's 0.73 and 0.72 — a 16.4% and 15.3% improvement.
  • Baseline progression: Base models score lowest (e.g., Human-QA: Gemini 1.0 0.34, GPT-4 0.38, Mistral Large 2 0.22, Llama 3.3 70B Instruct 0.35, Claude 3.5 Sonnet v2 0.42), Naive RAG improves on these (average improvement of 20-25% over base models), and specialized systems (Telco-RAG, Chat3GPP) improve further.
  • Ablation on Human-QA: Starting from a Naive RAG baseline (R 0.63, CR 0.6), each added component improved results: section-based chunking (0.72 / 0.73), query expansion (0.74 / 0.77), hierarchical retrieval (0.75 / 0.78), hybrid retrieval (0.78 / 0.79), reranking (0.79 / 0.8), filtering (0.82 / 0.81), and multi-modal fusion (0.87 / 0.83) for a total improvement of 38.1% in Recall and 38.3% in Claim Recall.
  • Model robustness: TelcoAI's Recall on Human-QA ranged from 0.83 to 0.87 across four generation models (Claude 3.5 Sonnet v2 0.87, Claude 3 Opus 0.85, Llama 3.3 70B Instruct 0.84, Mistral Large 2 0.83), a variation of only 2-4%, indicating the framework is not overly dependent on any specific generator.
  • TSpec-LLM benchmark: TelcoAI achieves 93% overall accuracy (Easy 93%, Intermediate 93%, Hard 92%), versus 92% for Chat3GPP with Claude 3.5 Sonnet v2. Base models ranged from 30% (Mistral Large 2) to 67% (Claude 3.5 Sonnet v2), and Naive RAG reached 73-82%.
  • Latency: TelcoAI's total latency is 11.81s (pre-retrieval 2.7s, retrieval 1.2s, post-retrieval 0.7s, generation 7.21s), compared with Base LLM 6.73s, Naive RAG 6.76s, Chat3GPP 8.07s, and Telco-RAG 9.91s. The paper characterizes this as a 3-5s additional latency over competing specialized systems that buys +16% Recall and +13.7% Claim Recall over Chat3GPP on Human-QA.
  • Hallucination: The conclusion reports high faithfulness (0.92) and minimal hallucination (0.05) on multi-document queries.

Methodology in Plain English

The system has two halves: ingestion and question answering.

On the ingestion side, 3GPP specifications in Microsoft Word (.docx) format are parsed and enriched. Structured metadata is extracted per document — release number (e.g., "R16"), series number (e.g., "23"), specification number (e.g., "23548-i30"), and version (e.g., "V18.3.0") — and stored alongside each chunk so retrieval can be filtered. Instead of fixed-size splitting, the document is decomposed recursively along its natural section and subsection headings, with each chunk mapped to its heading level and position within its parent section. For diagrams and tables, an LLM (Anthropic's Claude 3.5 Sonnet v2) generates a natural-language description that is concatenated into the chunk text next to a unique element marker, so visual content becomes retrievable by ordinary semantic matching while the original image remains linked.

On the question-answering side, four stages run in sequence. Pre-retrieval uses an LLM prompt to break a complex question into up to three standalone core sub-queries plus up to five follow-up queries, and separately extracts release, series, and specification metadata from the query. Context retrieval encodes each sub-query into a dense vector and performs a hybrid search that combines cosine similarity with BM25 lexical matching, weighted equally (alpha set to 0.5); the search works hierarchically, first matching smaller sub-section chunks and then expanding to their parent sections. Post-retrieval takes the top 40 matches, filters them by the extracted metadata, reranks the survivors, and keeps the top 5. Answer generation fuses chunks across sub-queries, pulls in the images referenced by the markers found in those chunks, and passes query, context, and images to Claude 3.5 Sonnet v2 with a prompt instructing it to reason across chunks and cite specifications and sections.

For evaluation, the authors use four datasets. Human-QA is their expert-curated set of 41 technical questions spanning 12 specifications from Releases 16-19, split into single-document retrieval (60%), cross-document synthesis (25%), and version comparisons (15%). Synthetic-QA uses the same corpus with questions generated by Claude 3.5 Sonnet v2 and validated with Mistral Large 2 and human experts. TSpec-LLM provides 30,137 documents spanning Releases 8-19 (535 million words) and an evaluation set of 100 multiple-choice questions by difficulty. Telco-DPR is named as a fourth dataset but no details about it are reported in the paper content. Human-QA and Synthetic-QA are scored with RAGChecker claim-level metrics (responses and ground truth decomposed into claims), with recall as the primary metric; TSpec-LLM is scored by accuracy across Overall, Easy, Intermediate, and Hard levels. Baselines include base LLMs (Gemini 1.0, GPT-4, Mistral Large 2, Llama 3.3 70B Instruct, Claude 3.5 Sonnet v2), Naive RAG, Telco-RAG, and Chat3GPP, each tested with multiple generation models. The implementation uses Claude 3.5 Sonnet v2 as the base LLM, Amazon Titan Text Embeddings v2 (1024 dimensions), temperature 0.7, top-5 chunks of 512 tokens with 20% overlap, BM25 via Elasticsearch with default parameters, and RAGChecker with Llama 3.3 70B Instruct as both claim extractor and checker.

Why This Matters

Impact on research. The paper argues that retrieval quality, not generation capacity, is the limiting factor for technical document QA in this domain: performance gains came from better context preparation (chunking, planning, filtering, multi-modal fusion) rather than longer generation, with generation times staying stable at 6-7s across all methods. It also shows that structure-aware ingestion and image captioning can be added to a general-purpose LLM without domain fine-tuning.

Real-world applications:

  • Engineers looking up how a specific procedure changed between 3GPP releases (the paper's example query concerns MEC interfaces across R17, R18, and R19 of specification 23.558).
  • Standards implementation teams needing to locate the sections and figures relevant to a feature across multiple specifications.
  • Telecom operators and vendors maintaining a searchable internal knowledge base over the specifications, as reflected in the Bouygues Telecom co-authorship.
  • QA over other long, hierarchical, diagram-heavy technical documentation where the same ingestion pattern could apply.

Industry relevance. The paper is a collaboration between the AWS Generative AI Innovation Center and Bouygues Telecom, framing the work as reducing cognitive load for engineers who must interpret, implement, and keep current with standards. The paper is explicit that the current implementation is a research prototype not yet optimized for system efficiency, inference latency, or integration overhead — the engineering work needed before production deployment.

Future Directions

  • File-format coverage: Extending multi-modal ingestion beyond .docx to PDFs with scanned images, Markdown with external media, and presentation slides.
  • Multi-turn conversation: Moving from single-turn queries to dialogue that maintains conversational state across complex, context-dependent technical support tasks.
  • Agentic planner and execution monitoring: Enabling the system to autonomously decide when to retrieve more context, cross-reference specifications, or request user clarification.
  • Efficiency and deployment: Optimizing system efficiency, inference latency, and integration overhead, and developing more sophisticated cross-version comparison capabilities and richer visual element handling.

Target Audience

Telecommunications engineers and standardization specialists who work with 3GPP specifications; RAG and applied LLM researchers interested in structure-aware chunking, agentic query planning, and multi-modal retrieval; and practitioners building domain-specific document QA systems over long, hierarchical, diagram-rich technical corpora. Readers should be comfortable with standard RAG terminology and evaluation metrics, as the paper assumes that background and does not define them from scratch.

Authors’ abstract

The 3rd Generation Partnership Project (3GPP) produces complex technical specifications essential to global telecommunications, yet their hierarchical structure, dense formatting, and multi-modal content make them difficult to process. While Large Language Models (LLMs) show promise, existing approaches fall short in handling complex queries, visual information, and document interdependencies. We present TelcoAI, an agentic, multi-modal Retrieval-Augmented Generation (RAG) system tailored for 3GPP documentation. TelcoAI introduces section-aware chunking, structured query planning, metadata-guided retrieval, and multi-modal fusion of text and diagrams. Evaluated on multiple benchmarks-including expert-curated queries-our system achieves $87\%$ recall, $83\%$ claim recall, and $92\%$ faithfulness, representing a $16\%$ improvement over state-of-the-art baselines. These results demonstrate the effectiveness of agentic and multi-modal reasoning in technical document understanding, advancing practical solutions for real-world telecommunications research and engineering.

Read the original paper