Research
How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
Overview Research area: Natural Language Processing; sociolinguistics of large language models; training-data auditing, tokenization fairness, and evaluation of regional language variation. Technical

- arXiv
- 2604.04204
- Published
- 2026-04-05
- Authors
- Mir Tafseer Nayeem, Davood Rafiei
AI summary
Overview
Research area: Natural Language Processing; sociolinguistics of large language models; training-data auditing, tokenization fairness, and evaluation of regional language variation.
Technical level: Intermediate. The conceptual argument is accessible, but the paper relies on familiarity with pretraining corpora, instruction tuning/RLHF data, subword tokenization (fertility, vocabulary size), and per-token negative log-likelihood.
Scope (one sentence): The paper traces a structural advantage for American English (AmE) relative to a controlled British English reference, Ref(BrE), across three stages of the LLM pipeline — data exposure, representation, and generation — using a curated 1,813-pair variant resource and a training-free alignment method called DiAlign.
What This Paper Is About
Widely used LLM platforms expose "English (US)" as a primary English setting even though English is globally diverse, so the authors ask how AmE becomes the implicit default. They treat this not as a stylistic preference but as structural bias: a systematic linguistic advantage that should be visible in several places at once, from the corpora used in training, through how tokenizers and models represent regional word forms, to what models actually generate. The goal is to localize where the asymmetry appears, whether it persists across pipeline components, and therefore where component-level fixes could be aimed.
Key Contributions
- A curated matched-variant resource: 1,813 manually compiled AmE–Ref(BrE) pairs covering orthographic and lexical differences, drawn from linguistic references and web-based lexicons.
- DiAlign, a dynamic training-free alignment estimator: A method that decomposes text into overlapping 2- to 5-grams, looks up separate AmE and BrE reference frequencies in Google Books Ngrams, computes signed log-frequency ratios, weights spans (extra weight for curated lexicon matches), and normalizes positive and negative evidence into (P_AmE, P_Ref(BrE)).
- A pipeline-wide triangulation: Joint auditing of data exposure (6 pretraining corpora, 21 public post-training datasets), representation (9 tokenizers, tokenizer provenance, 10 model checkpoints for prediction cost), and generation (10 models on two datasets under two prompting conditions).
- A cited first: The authors describe this as the first rigorous pipeline-wide experimental study of structural bias across major phases of LLM development, plus practical recommendations for corpus construction/filtering, tokenizer design and provenance auditing, post-training data selection, and evaluation.
Main Findings
- AmE is favored in all six pretraining corpora. Book Corpus (2015), Wikipedia (2024), Common Crawl C4 (2020), Falcon RefinedWeb (2023), RedPajama (2024), and Dolma (2024) all lean toward AmE, and all six differ significantly from parity (Wilcoxon signed-rank, p < 0.01). Orthographic variants show the strongest skew, with AmE spellings often exceeding 70% of observed pair frequencies; vocabulary contrasts are more balanced but still lean AmE. DiAlign shows the same direction for grammatical, structural, stylistic, and multi-word evidence. Example: Falcon RefinedWeb 77.34% AmE vs 22.66% Ref(BrE) orthographic, 68.35% vs 31.65% vocabulary, 74.04% vs 25.96% DiAlign.
- DiAlign is validated at 93.18% accuracy. On a balanced benchmark of 1,500 news passages (750 from HuffPost U.S. News, 750 from BBC England) using Google Books frequencies from 1950–2022, DiAlign reaches 90.67% precision, 96.25% recall, and 93.38 F1. Restricting the reference window to 2000–2022 gives 92.91% accuracy and 93.12 F1. Errors concentrate near the decision boundary.
- AmE is favored in all 21 post-training datasets. Across 11,574,254 samples there are 28,860,119 explicit variant occurrences, of which 76.17% are AmE. Instruction tuning/SFT is 75.95% AmE; task mixture 71.77%; chat 84.35%; human feedback 83.08%; reward-model feedback 84.72%; safety preference 86.86%. Preference, reward-modeling, and safety resources together contain 4,666,867 occurrences and are 80.94% AmE. Dataset-level values range from 65.29% (Databricks Dolly 15k) to 89.72% (WizardLM Evol-Instruct v2).
- AmE forms are generally tokenized more compactly. Across the evaluated tokenizers, Ref(BrE) variants need more subword tokens. The largest gap is Δv = 18.72% for vocabulary contrasts; orthographic gaps among AmE-favoring tokenizers range from 2.85% to 5.42%. Ref(BrE) forms appear more often in the 3+ subword category. Differences are statistically significant (Wilcoxon, p < 0.01) except Velvet-2B in the orthographic category, where BrE is slightly favored (Δo = −0.69%). DeepSeek-V3 has the smallest vocabulary gap (Δv = 12.66%), and Gemma-3-27B has the lowest fertility for both forms, consistent with its 262K-token vocabulary.
- Tokenizer provenance can transmit representational behavior. Of nine tokenizers, seven use independent or source vocabularies, while Llama-3.3 and StableLM-2 inherit or extend U.S.-origin tokenizer artifacts. StableLM-2, associated with a UK-developed model, shows 100% token-boundary identity with GPT-4 on both variant sets, with all 100,256 shared base tokens retaining identical rank/id assignments.
- Matched BrE forms carry higher prediction cost. Paired counterfactual sentences differing only in the target regional form are verified as semantically equivalent (closer in contextual representation space than character-perturbed and unrelated controls). Yet all ten evaluated checkpoints (five base/post-trained model pairs) assign higher per-token loss to Ref(BrE) on the strict equal-token-count slice, with all paired-bootstrap 95% confidence intervals excluding zero. The gap is larger after post-training in every family: Gemma, Ministral, Llama, StableLM, and Falcon3.
- AmE is the dominant generation default under neutral English prompting. On Natural Questions (formal) and ELI5 (informal), most models produce majority AmE-classified outputs under default English. Examples: GPT-4o 79.00% AmE on NQ and 77.00% on ELI5; Gemini-2.0-flash 76.00% and 75.33%; Gemma-3-27B 69.33% and 68.33%.
- British-English prompting shifts but does not consistently eliminate the default. Requesting British English (en-GB) moves outputs toward Ref(BrE), but substantial AmE alignment often remains, especially on the formal NQ setting. GPT-4o drops to 45.33% on NQ and 34.67% on ELI5; Llama-3.3-70B falls to about 30.00% AmE-classified outputs on ELI5. StableLM-2 (74.00% to 69.67% on NQ) and Velvet-2B (72.91% to 69.33% on NQ) shift far less. Alignment is effectively uncorrelated with response length (r = 0.017).
Methodology in Plain English
The authors pick British English as a fixed yardstick, Ref(BrE), and ask at each pipeline stage whether AmE gets an advantage.
First they build a resource of 1,813 matched word pairs that differ between the two varieties (spelling contrasts like -ize/-ise and -og/-ogue, plus vocabulary contrasts). For long documents, they use DiAlign: the text is chopped into every overlapping 2- to 5-word span, named-entity and stopword-only spans are discarded, and each remaining span is looked up in separate American and British Google Books Ngram frequency series. A span's signed log-frequency ratio indicates which variety it favors, weighted down when the two frequencies are close and weighted up when the span contains a curated variant. Positive and negative evidence are summed and normalized into AmE and Ref(BrE) alignment probabilities.
For data exposure, they count matched variants in six pretraining corpora (computing pair-level AmE probabilities from corpus frequencies) and in 21 public post-training datasets spanning instruction tuning, chat, human feedback, preference data, reward modeling, and safety alignment. For representation, they measure tokenizer fertility (subword tokens per word) and granularity (how often a form takes 1, 2, or 3+ subwords), then check tokenizer lineage by vocabulary overlap and token-boundary identity. Separately, they build paired counterfactual sentences that differ only in the regional form, verify semantic equivalence via contextual embeddings against controls, and compare per-token negative log-likelihood — restricted to pairs with identical sentence token counts so the loss gap cannot be explained by unequal tokenization. For generation, they prompt ten models on Natural Questions (formal) and ELI5 (informal), after filtering out questions containing lexicon variants to avoid priming, and compare default English against British English (en-GB) requests for 50-word answers, classifying each response with DiAlign.
Regarding the appendices (construction and filtering details, ablations, per-model bootstrap intervals, prompt figures, and the ethics/limitations discussion), the supplied text is truncated at Appendix C; those details are referenced but their contents are not reported in the available content.
Why This Matters
The paper argues that a pipeline-wide AmE advantage is not merely a stylistic matter. Because LLMs are embedded in educational, professional, legal, and public-sector infrastructure — AI is used in at least one area of government in 35 of 36 OECD countries — their defaults operate at global scale, and the authors link the pattern to linguistic homogenization, epistemic injustice, and inequity in global AI deployment.
Real-world applications:
- Corpus curation and filtering: knowing that orthographic contrasts skew hardest in pretraining data (often above 70% AmE) tells data teams where rebalancing or provenance-aware sampling would matter most.
- Tokenizer design and audits: fertility and granularity differences, plus tokenizer reuse across model families, give concrete metrics and a lineage check for teams designing or inheriting vocabularies.
- Post-training data selection: the skew appears in preference, reward-modeling, and safety resources (80.94% AmE across 4,666,867 occurrences), which are exactly the datasets that shape model behavior after pretraining.
- Evaluation and deployment for non-US English users: the generation results quantify how much regional steering a prompt provides and where it falls short, which matters for public-sector, educational, and legal deployments serving BrE-influenced users.
Industry relevance: Tokenizer provenance results suggest that model developers in one country can inherit representational conventions from another through reused vocabularies, and the finding that the loss gap widens after post-training in every evaluated model family points to the alignment stage, not just the pretraining crawl, as a place where the asymmetry is reinforced.
Future Directions
- Controlled intervention studies: the authors note that isolating a causal post-training contribution to the loss gap requires controlled intervention, which they have not run.
- Extending beyond the binary AmE–Ref(BrE) contrast: the authors state that DiAlign generalizes to other variety pairs with suitable reference distributions and linguistic resources, and flag scope, limitations, and risks of binary regional comparison as a topic for discussion.
- Component-level improvement and re-auditing: the paper motivates practical recommendations for corpus construction, tokenizer design and provenance auditing, post-training data selection, and evaluation, but does not report results from applying them.
- Explaining why prompting falls short: British-English prompting shifts output without consistently removing the default, and the shift's magnitude varies by model and register; the mechanism behind that variation is left open.
Target Audience
Researchers and practitioners working on multilingual and regional language equity, training-data auditing, tokenization fairness, and LLM evaluation will get the most from this paper, as will NLP ethics and sociolinguistics researchers interested in how standardization and digital dominance show up in model infrastructure. Data engineers and model developers selecting corpora, vocabularies, or post-training data can use the results operationally. Readers should be comfortable with tokenizer statistics and language-model loss; readers seeking full methodological detail will need the appendices, which are truncated in the supplied content and are referenced throughout rather than fully reproduced.
Authors’ abstract
Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)" become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.