Research
Detecting Legend Items on Historical Maps Using GPT-4o with In-Context Learning
Overview Research area: Computer vision and geospatial information retrieval, specifically document layout analysis and multimodal large language model prompting applied to the digitization of histori

- arXiv
- 2510.08385
- Published
- 2025-10-09
- Authors
- Sofia Kirsanova, Yao-Yi Chiang, Weiwei Duan
AI summary
Overview
Research area: Computer vision and geospatial information retrieval, specifically document layout analysis and multimodal large language model prompting applied to the digitization of historical geological maps.
Technical level: Intermediate. The paper assumes familiarity with object detection metrics (IoU, F1) and large language model prompting, but the pipeline itself is conceptually simple.
Scope: The paper presents a training-free pipeline that segments legend areas from scanned USGS geological maps with LayoutLMv3 and then uses GPT-4o with in-context learning and structured JSON prompts to predict bounding boxes for legend items and their linked descriptions.
What This Paper Is About
Historical geological maps encode information about rock units and mineral deposits through symbols that only make sense if you can read the map legend. Legends vary wildly in layout across map series and time periods, and there is no large annotated dataset of them, so supervised model training is difficult. This paper's goal is to automatically detect pairs of legend symbols and their text descriptions in a structured, machine-readable form.
Key Contributions
- A training-free, in-context learning (ICL) approach using GPT-4o to detect and link legend items and descriptions on historical maps, without additional training or large-scale annotations.
- A structured JSON prompting format that produces machine-readable bounding box coordinates, designed to integrate with geospatial search and digitization pipelines.
- An empirical analysis of prompt design, varying the number of in-context examples (5, 10, 15, and 20) to identify the optimal configuration.
- An extension of the DIGMAPPER project's legend detection module, pairing a layout-aware segmentation step with the GPT-4o detection step.
Main Findings
- Headline performance: GPT-4o with structured JSON prompts outperforms the baseline, achieving 88% F-1 and 85% IoU (reported in the abstract; the same values appear in Table 2 for legend items).
- Best in-context example count is 15. With 15 examples, legend items scored IoU 0.85 and F1 0.88, and descriptions scored IoU 0.84 and F1 0.92.
- Performance rises then falls with more examples. At 5 examples, legend items scored IoU 0.78 / F1 0.82 and descriptions IoU 0.76 / F1 0.84. At 10 examples these rose to IoU 0.82 / F1 0.86 (items) and IoU 0.81 / F1 0.88 (descriptions). At 20 examples performance declined to IoU 0.83 / F1 0.86 (items) and IoU 0.82 / F1 0.88 (descriptions), which the authors attribute to overly long prompts introducing noise.
- The GPT-4o approach beats the LayoutLMv3 baseline. LayoutLMv3 alone scored IoU 0.69 / F1 0.72 for legend items and IoU 0.75 / F1 0.79 for descriptions, versus IoU 0.85 / F1 0.88 and IoU 0.84 / F1 0.92 for GPT-4o with 15 examples.
- The dominant failure mode is layout sensitivity. In a failure case with a dense multi-column legend, several descriptions were missed or merged into oversized bounding boxes, and some legend items were paired incorrectly.
- Evaluation setup details: 40 annotated maps randomly selected from the DARPA–USGS historical map dataset, all high-resolution .tiff scans; a prediction counts as correct at an IoU threshold of 0.5 with ground truth boxes.
- Implementation details: all experiments used GPT-4o (April 2025 version) through the OpenAI API, with a deterministic pipeline, fixed prompts, and temperature = 0. Each query averaged 1.2k input tokens and 0.6k output tokens.
Methodology in Plain English
The pipeline has two steps. First, a fine-tuned LayoutLMv3 model called LARA classifies regions of the scanned map and identifies the block containing the legend, which is then cropped out as an image. This matters because legends appear on the right side, at the bottom, or in irregular multi-column arrangements depending on the map.
Second, GPT-4o is asked to produce bounding boxes for legend items and their matching descriptions in that cropped legend. The model is not retrained; instead it is given a prompt built from three parts: a cropped legend area from an already-annotated map, a JSON text block defining the task and listing annotated item-description bounding box pairs from that example, and the cropped legend from a new unseen map. The JSON lists 15 example pairs and a placeholder entry for the target map with coordinates set to "??". GPT-4o returns a JSON object in which the placeholders are replaced with numeric coordinates for the detected pairs.
Why This Matters
Impact on research: The work shows that a general-purpose multimodal LLM can perform structured document-layout extraction without any fine-tuning or large annotated dataset, which reframes legend parsing as a few-shot prompting problem rather than a supervised learning one. It also quantifies how example count and prompt design affect accuracy, and it documents layout sensitivity as the central open problem.
Real-world applications:
- Library and archive search: once legends are digitized, scanned map collections become queryable by geological unit names, symbols, or color patterns instead of remaining unsearchable images.
- Geological and survey work: users can find all maps in a collection containing a specific rock unit or fault symbol.
- Color-pattern retrieval: dominant fill colors from each legend item can be stored so users can search for maps with similar rock-unit shades.
- Digitization pipelines: the structured bounding-box output feeds downstream georeferencing, feature extraction, and multimodal search tools.
Industry relevance: The method is part of DIGMAPPER, a modular automated map digitization system developed under the DARPA CriticalMAAS program, and the work was supported in part by DARPA (Agreement No. HR00112390132 and Contract No. 140D0423C0093). Organizations scanning thousands of historical maps at scale are the direct beneficiaries of a training-free parsing step that avoids the cost of building large annotated legend datasets.
Future Directions
- Improving robustness on tightly packed, multi-column legends, where descriptions are missed, merged into oversized boxes, or paired incorrectly.
- Better handling of irregular or custom legend symbols, which the paper names alongside multi-column layouts as a remaining limitation.
- Refining prompt design and example selection, since the paper identifies these as crucial factors and finds that the optimal number of examples sits at 15 rather than at either extreme tested.
- Extending evaluation beyond the 40 annotated USGS maps used here, given that the dataset covers a range of USGS styles, years, scales, and layout formats but the paper does not report results on non-USGS map series.
Target Audience
Researchers and practitioners working on historical map digitization, geospatial information retrieval, and document layout analysis, particularly those interested in applying multimodal LLMs to structured extraction tasks. It is also relevant to digital library and archive engineers building searchable map collections, and to readers following the DIGMAPPER and DARPA CriticalMAAS efforts. The paper is most useful to someone who already understands object detection metrics and is weighing whether a training-free prompting approach can replace a supervised model for their own layout extraction problem.
Authors’ abstract
Historical map legends are critical for interpreting cartographic symbols. However, their inconsistent layouts and unstructured formats make automatic extraction challenging. Prior work focuses primarily on segmentation or general optical character recognition (OCR), with few methods effectively matching legend symbols to their corresponding descriptions in a structured manner. We present a method that combines LayoutLMv3 for layout detection with GPT-4o using in-context learning to detect and link legend items and their descriptions via bounding box predictions. Our experiments show that GPT-4 with structured JSON prompts outperforms the baseline, achieving 88% F-1 and 85% IoU, and reveal how prompt design, example counts, and layout alignment affect performance. This approach supports scalable, layout-aware legend parsing and improves the indexing and searchability of historical maps across various visual styles.