Skip to content
AI.info

Research

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

Overview Research area: Document understanding and computer vision — specifically vision-language models (VLMs) applied to document OCR combined with schema-guided structured JSON extraction. Technica

arXiv
2610.11818
Published
2026-10-08
Authors
Uddipan Basu Bir, Vincent Christlein, Andreas Maier, Mathias Zinnen

AI summary

Overview

Research area: Document understanding and computer vision — specifically vision-language models (VLMs) applied to document OCR combined with schema-guided structured JSON extraction.

Technical level: Intermediate. The paper assumes familiarity with VLMs, fine-tuning concepts such as LoRA/QLoRA, and OCR evaluation metrics (CER, ANLS*, mAP-F1), but each method is described concretely enough for a practitioner to follow.

Scope (one sentence): A controlled comparative benchmark of eight open-source lightweight VLMs (≤ 7B parameters) on three German university heritage collections, measuring how well they turn document images directly into schema-compliant JSON, plus three follow-up studies on hyperparameter optimization, classical image preprocessing, and multi-stage training.

What This Paper Is About

Large closed-source VLMs set strong benchmarks for document understanding, but institutional archives cannot easily adopt them because of data-autonomy concerns, recurring API costs, and the environmental footprint of hyperscale computing. This is particularly acute in heritage digitization, where records contain historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts that must be mapped into highly specific database schemas. The paper asks whether small, locally deployable open-source VLMs can perform this "OCR-to-structure" task reliably, producing not just readable text but valid JSON that follows a fixed archival schema.

Key Contributions

  1. A benchmark of eight lightweight VLMs (≤ 7B parameters) on three specialized archival collections (Erlangen-Prehistoric, Pforzheim Jewelry, Regensburg Scalalogy), evaluating structured OCR-to-JSON as a private, sustainable alternative to commercial APIs.
  2. A constraint-aware evaluation protocol that accounts for differing model interfaces — some models accept only fixed OCR-style prompts or task tokens (e.g., <s_cord-v2>, <OCR>), while others accept custom natural-language prompts requesting structured JSON.
  3. Independent ablation studies on the top-3 models, quantifying the separate effects of (i) hyperparameter optimization via Optuna, (ii) classical image preprocessing (illumination flattening, denoising, CLAHE, letterbox resize), and (iii) multi-stage curriculum training.
  4. An analysis of single-dataset versus multi-dataset fine-tuning, where one joint checkpoint is trained on the union of all training sets and compared against per-dataset checkpoints, revealing both cross-collection gains and negative transfer.

Main Findings

  • Instruction-tuned general-purpose VLMs win: Across zero-shot, few-shot, and fine-tuning, the strongest models are the instruction-tuned general-purpose VLMs — Qwen2.5-VL, Phi-3.5-Vision, and Gemma-3 — which outperform several OCR-specialized alternatives in structured generation. These three were selected as the top-3 for further study.

  • JSON validity separates the field: JSON validity stays in the range of approximately 95%–100% for Gemma-3, Phi-3.5-Vision, and Qwen2.5-VL across settings, but remains much lower for Donut, Florence-2, and PaddleOCR-VL even after fine-tuning. Donut is the clearest example: it can achieve lower CER / higher ANLS* after fine-tuning yet still yields 0 mAP-F1 because its outputs often fail strict JSON parsing.

  • The three metrics measure different things: CER and ANLS* fall back to raw-text comparison when JSON parsing fails, whereas mAP-F1 requires valid JSON and is set to 0 otherwise. A model can therefore post low CER and high ANLS* while still scoring low mAP-F1 if values are assigned to the wrong keys, fields are missing or merged, or extra fields are hallucinated.

  • Zero-shot performance is weak on the hardest schemas: Zero-shot results are weak across the board on Regensburg Scalalogy, the high-complexity nested schema — e.g., Donut-Base scores 90.1 CER with 0.0 ANLS* and 0.0 mAP-F1, and Florence-2-Large scores 75.6 CER with 0.0 ANLS* and 0.0 mAP-F1. On Erlangen-Prehistoric, Florence-2-Large records 45.6 CER, 50.4 ANLS*, and 0.0 mAP-F1, and Qwen2.5-VL-7B records 21.4 CER, 54.0 ANLS*, and 56.2 mAP-F1.

  • Fine-tuning transforms the best models: Qwen2.5-VL-7B reaches 5.8 CER, 91.0 ANLS*, and 90.0 mAP-F1 on Erlangen-Prehistoric, 0.5 CER, 99.2 ANLS*, and 99.4 mAP-F1 on Pforzheim Jewelry, and 7.6 CER, 87.0 ANLS*, and 62.8 mAP-F1 on Regensburg Scalalogy. Gemma-3-4B reaches 10.4 CER, 88.6 ANLS*, 83.1 mAP-F1 on Erlangen; Phi-3.5-Vision reaches 17.9 CER, 86.5 ANLS*, 82.2 mAP-F1 on Erlangen.

  • Few-shot helps most where zero-shot was weakest: With k = 1 example for Pforzheim Jewelry and k = 2 for Erlangen-Prehistoric and Regensburg Scalalogy, Qwen2.5-VL-7B improves on Erlangen-Prehistoric from 21.4 CER / 54.0 ANLS* / 56.2 mAP-F1 in zero-shot to 18.9 / 79.0 / 74.7.

  • Hyperparameter optimization helps unevenly: The largest gains appear on Regensburg Scalalogy for all three models. Gemma-3-4B improves there from 14.6 CER / 73.8 ANLS* / 46.9 mAP-F1 (default) to 3.4 / 92.3 / 77.9 (HPO), and Qwen2.5-VL-7B improves from 7.6 / 87.0 / 62.8 to 3.1 / 89.4 / 67.1. On Erlangen-Prehistoric, HPO clearly benefits Qwen2.5-VL-7B (5.8 → 4.3 CER; 91.0 → 93.4 ANLS*; 90.0 → 91.2 mAP-F1) but is mixed for Gemma-3-4B and Phi-3.5-Vision — Gemma's Erlangen mAP-F1 actually drops from 83.1 to 65.4.

  • Classical preprocessing is mostly mixed and often harmful: It improves CER for Phi-3.5-Vision on Regensburg Scalalogy (6.3% → 5.5%) and for Qwen2.5-VL-7B on Erlangen-Prehistoric (5.8% → 5.2%), but these gains do not consistently translate into better structured extraction — Phi-3.5-Vision's Regensburg mAP-F1 drops from 72.0 to 61.1 despite the CER gain. Preprocessing is clearly harmful in several settings: Erlangen-Prehistoric for Phi-3.5-Vision, Pforzheim Jewelry for Gemma-3-4B and Qwen2.5-VL-7B, and Regensburg Scalalogy for Gemma-3-4B and Qwen2.5-VL-7B.

  • Multi-stage training is not a guaranteed improvement: It consistently benefits Gemma-3-4B on all three datasets (Erlangen 10.4 → 8.1 CER, 88.6 → 89.9 ANLS*, 83.1 → 86.5 mAP-F1; Regensburg 14.6 → 10.3 CER, 73.8 → 80.2 ANLS*, 46.9 → 63.8 mAP-F1). For Phi-3.5-Vision it notably improves Regensburg Scalalogy (6.3 → 4.9 CER; 72.0 → 72.5 mAP-F1) while other changes are small and sometimes negative. For Qwen2.5-VL-7B it slightly helps Erlangen but degrades Pforzheim Jewelry (ANLS* 99.2 → 88.5, mAP-F1 99.4 → 90.2) and Regensburg Scalalogy.

  • Multi-dataset training is model-dependent, not a uniform trade-off: Gemma-3-4B's joint checkpoint beats its single-dataset checkpoints on all three datasets (Erlangen 10.4 → 8.7 CER and 83.1 → 87.9 mAP-F1; Pforzheim 0.9 → 0.3 CER; Regensburg 14.6 → 9.1 CER, 73.8 → 81.5 ANLS*, 46.9 → 68.0 mAP-F1). Qwen2.5-VL-7B shows the opposite: the multi-dataset checkpoint degrades CER and ANLS*

Authors’ abstract

While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.

Read the original paper