Skip to content
AI.info

Research

MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering

Overview Research area: Natural language processing, specifically locale-aware / geographically grounded question answering and large language model pretraining. Technical level: Intermediate. The pap

MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering
arXiv
2601.15236
Published
2026-01-21
Authors
Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos

AI summary

Overview

Research area: Natural language processing, specifically locale-aware / geographically grounded question answering and large language model pretraining.

Technical level: Intermediate. The paper assumes familiarity with decoder-only transformer pretraining, perplexity, token budgets, and multiple-choice evaluation, but its central idea (keep geographic metadata attached to documents during pretraining) is easy to grasp.

Scope: The paper introduces LocalNewsQA, an 18,700-item English-news benchmark for measuring whether a model changes its answer when the locale changes, and MAPLE, a controlled family of decoder-only models pretrained with document-level geographic metadata (source URL, country, continent), comparing them against metadata-free controls trained on identical data with the same token budget, architecture, and optimization.

What This Paper Is About

Large language models can memorize locale-specific facts from many countries but often fail to select the right one when the locale changes, defaulting to a single globally dominant answer. The authors formalize this as localized knowledge disambiguation and ask whether geographic metadata that already sits inside a training corpus (source URL, country of origin, continent) can be preserved during pretraining so that one global model learns to condition its factual predictions on locale. To test this they build a paired benchmark, LocalNewsQA, that scores whether the same model actually switches answers when only the locale changes.

Key Contributions

  1. Task and benchmark. The authors define localized knowledge disambiguation and release LocalNewsQA, an 18,700-item English-news benchmark that pairs the same question across two locales with different gold answers, making paired locale switching directly observable. It covers 17 countries across 4 continents and 8 topic families.
  2. Controlled pretraining. They introduce MAPLE, a family of decoder-only models trained from scratch with document-level geographic metadata from the NOW corpus, tested while holding data, architecture, and token budget fixed against metadata-free controls.
  3. Efficient localization. With inference-time metadata held fixed, metadata-conditioned pretraining improves locale switching, and the gains increase from 1B to 3B parameters.
  4. Mechanism and transfer. Ablations, comparisons crossing training-time and inference-time metadata, and external benchmark evaluations suggest the full effect requires geographic provenance during both pretraining and inference, and that it transfers to held-out evaluations.

Main Findings

  • Paired switching improves with metadata pretraining, and more so at scale. Holding inference-time locale information fixed, the paired margin-switch rate on the 1,700 ambiguous LocalNewsQA pairs rises from 6.04% to 7.78% at 1B (+1.74 points) and from 21.36% to 24.13% at 3B (+2.77 points).
  • Target-locale accuracy also improves. Comparing the (T+, I+) setting against the (T-, I-) negative control on the ambiguous split yields gains of +6.7 points at 1B and +9.2 points at 3B; the explicit split improves too, but less sharply because those questions already name the locale. Bootstrap intervals exclude zero at both sizes.
  • The negative control behaves as expected. Both I- settings produce zero paired switches by construction, because the target and contrast prompts are identical.
  • Metadata lowers region-specific perplexity. Models trained and evaluated with full metadata achieve the lowest perplexity on every same-continent test set, with crossed settings (metadata at only one stage) falling in between. Same-continent perplexity is consistently lower than cross-continent perplexity, so local models remain local.
  • One global model can serve multiple regions. A single metadata-conditioned global model at 1B lowers perplexity relative to the metadata-free control on every continent-specific test set and on the combined global test set, approaching local-model quality from 500M to 1B while remaining stronger on the combined set. The pattern continues at 3B.
  • Metadata speeds convergence. Metadata-conditioned models reach any given perplexity target using fewer training tokens, an efficiency benefit that persists across model sizes.
  • The URL alone carries most of the signal. Among five 1B metadata-granularity variants (URL, URL + Country, URL + Continent, Country, Continent), URL-only nearly matches full metadata and substantially outperforms the no-metadata baseline. Country and Continent trail URL and add little when paired with it.
  • Metadata cannot substitute for missing data. In leave-one-continent controls, excluding any continent raises held-out and global perplexity regardless of metadata, so provenance organizes evidence the model has seen but does not replace regional coverage it lacks.
  • Transfer to external benchmarks. On six external benchmarks (GeoMLaMA, GlobalOpinionQA, WorldValuesBench, NormAD, BLEnD, Global-MMLU-CS), MAPLE remains competitive with size-matched models. At 3B it shows statistically significant gains on GlobalOpinionQA (+8.4), NormAD (+5.1), and BLEnD (+13.5), with positive point estimates on the remaining three tasks.
  • The effect is not decoder-specific. In an architecture sanity check, a matched 1B Qwen2-style pair was also trained. The metadata-trained checkpoint scored 9.33 [9.00, 9.63] on the metadata test versus 10.45 [10.09, 10.79] for the metadata-free checkpoint; on the no-metadata test the metadata-trained checkpoint scored 10.45 [10.09, 10.79] against 9.85 [9.51, 10.16] for the metadata-free checkpoint.
  • Benchmark quality was audited. In a three-annotator audit of a stratified 100-item gold sample, majority vote accepted 96/100 items, including 50/50 explicit and 46/50 ambiguous, with 98.0% average pairwise agreement and Fleiss' kappa = 0.76 on the binary accept/reject decision.

Methodology in Plain English

The authors take the NOW news corpus, which already attaches a publication year, source URL, and country of origin to each document, and map countries to continents using the United Nations geoscheme. They then compare two ways of feeding the same text to a model: one that prepends structured metadata (URL, country, continent) before the article title and content, and one that omits those fields entirely. Everything else is held fixed, including the model family, tokenizer, optimizer, token budget, sequence length, and random seed, so any behavioral difference can be attributed to the metadata intervention.

They train 33 models from scratch at two main sizes. At 500M, ten models cover four continents times two metadata settings, plus two global settings. At 1B, 23 models cover those same ten settings, plus five provenance-granularity ablations and eight leave-one-continent controls. They additionally train two 3B global models to test scaling. All local and global models are trained on 41.9B tokens; local models sample from a single continent, while global models mix all continents before sampling. Each continent has a held-out test set of 1,000 documents, and the global test set has 1,000 documents with 250 sampled uniformly from each continent. Perplexity is computed over non-metadata tokens only, so improvements reflect better content prediction rather than metadata memorization.

For the benchmark, candidate question frames are generated, then filtered so that only rows passing split-specific source and formatting checks survive. The explicit split names the locale in the question and tests localized factual recall; the ambiguous split omits the locale from the question text and supplies it through the evaluation context, with a paired contrast locale that has a different gold answer. Every question includes source evidence for auditability.

Evaluation uses two paired metrics. Exact pair is the stringent case where the model must be correct on both the target and contrast locales. Margin switch is softer, requiring only that the model's score ordering prefer the locale-appropriate answer on both sides. MAPLE checkpoints are scored using length-normalized log-probability over each option's token sequence, with the crossed paired-switch comparison averaged over five answer-order seeds.

Why This Matters

Impact on research. The work isolates a behavior that accuracy-only geographic and cultural evaluations cannot measure: whether a model actually changes its answer when the locale changes. By holding language fixed in English throughout, it separates locale-conditioned answer selection from translation quality and multilingual transfer, which confounds benchmarks that vary language alongside locale. It also extends a line of work on metadata conditioning, moving beyond perplexity as the outcome of interest to a directly measurable behavioral change at inference time.

Real-world applications.

  • Locale-aware assistants and search systems that should return country-specific answers about institutions, school calendars, sports leagues, or public agencies rather than a single globally dominant default.
  • News and media products that surface locally valid facts and references tied to the outlet's country of origin.
  • Retrieval and question answering pipelines that need to route or rank competing facts by geography without retraining separate models per region.
  • Data curation tooling: the finding that source URLs carry most of the geographic signal argues for preserving provenance rather than stripping corpora to plain text during preprocessing.

Industry relevance. The result that a single global model can retain multiple regional distributions without region-specific parameters is directly relevant to practitioners who cannot afford per-region models. The token-efficiency finding, that metadata-conditioned models reach a given perplexity with fewer training tokens, is a practical cost argument. The finding that metadata cannot compensate for regions absent from training also tells teams where to spend effort: metadata conditioning and broader geographic data collection serve different purposes and neither replaces the other.

Future Directions

  1. Scale beyond 3B. The from-scratch experiments cover 500M to 3B parameters; the authors did not test 70B or larger because that scale was infeasible under their academic compute budget, so whether paired-switching gains persist beyond 3B remains untested.
  2. Broader transfer through the same factorial design. Additional model families, multilingual corpora, open-web sources such as Common Crawl, and non-news domains where provenance may encode different kinds of context.
  3. Extending the protocol beyond geography. The paired-switching design is not inherently geographic. Temporal provenance (publication date) could distinguish "Who is the president?" across time periods, and source-type provenance (books versus web data) could help weight formal versus colloquial usage.
  4. Separating evaluation uncertainty from pretraining stochasticity. Repeated pretraining runs, and richer benchmark work combining the controlled template with naturally occurring user questions and richer human annotation.

Target Audience

Researchers working on geographic and cultural language model evaluation, locale-aware question answering, and data-centric pretraining. It is also useful for practitioners who build multilingual or regionally deployed systems and need to decide whether to preserve provenance metadata in their corpora, and for benchmark designers interested in paired or contrastive evaluation designs rather than single-locale accuracy. Readers should be comfortable with pretraining concepts such as token budgets, perplexity, and multiple-choice scoring.

Authors’ abstract

Large language models can memorize competing locale-specific facts yet fail to select among them when the locale changes, defaulting instead to a single globally dominant answer. We formalize this as localized knowledge disambiguation and introduce LocalNewsQA, an 18,700-item English-news benchmark that pairs the same question across two locales and scores whether a model actually switches its answer when the locale changes. We also introduce MAPLE, a controlled family of decoder-only models pretrained with document-level geographic metadata (source URL, country, and continent) already present in the training corpus, and compare it to metadata-free controls trained on identical data with the same token budget, architecture, and optimization. In controlled experiments at 1B and 3B, with inference-time metadata fixed, pretraining with metadata in MAPLE produces measurable switching and improves accuracy on questions whose correct answer depends on locale. Ablations and external-benchmark evaluations further suggest that locale-conditioned prediction benefits from geographic provenance learned during pretraining and that these benefits strengthen at larger model sizes.

Read the original paper