Research
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
Overview Research area: Natural Language Processing — text normalization (TN) as a preprocessing step for Text-to-Speech (TTS) systems, using few-shot prompting with Large Language Models. Technical l

- arXiv
- 2511.03080
- Published
- 2025-11-05
- Authors
- Michel Wong, Ali Alshehri, Sophia Kao, Haotian He
AI summary
Overview
Research area: Natural Language Processing — text normalization (TN) as a preprocessing step for Text-to-Speech (TTS) systems, using few-shot prompting with Large Language Models.
Technical level: Intermediate. The problem domain (converting written forms like "17°" or "CHF 20" into spoken equivalents) is intuitive, but the paper assumes familiarity with concepts such as word error rate, BLEU, in-context learning, and weighted finite-state transducers.
Scope: The paper proposes PolyNorm, an LLM prompting framework for multilingual text normalization, and evaluates it against a production rule-based system across eight languages using a newly released 27-category benchmark.
What This Paper Is About
Text normalization converts written text dense with numbers, abbreviations, dates, and symbols into the words a speaker would actually say — so "4/18" becomes "april eighteenth" and "CHF 20" becomes the spoken currency form. Traditional systems built on handcrafted rules and weighted finite-state transducers achieve high accuracy but are expensive to build and maintain, and they scale poorly to low-resource or morphologically rich languages. PolyNorm replaces that rule engineering with step-by-step prompts and a small set of curated in-context examples, aiming to make normalization cheaper, faster to iterate on, and portable across languages.
Key Contributions
-
PolyNorm, a prompt-based TN framework. A language-agnostic instruction prompt combined with localized in-context learning (ICL) examples, applied across eight languages without fine-tuning or rule authoring.
-
PolyNorm-Benchmark. A multilingual dataset spanning 27 normalization categories with 20 examples per category, for a total of 540 high-quality data points per language, generated with DeepSeek-R1 and then edited and verified by internal language experts. It is released at
https://github.com/apple/ml-speech-polynorm-bench. -
A language-agnostic pipeline for automatic data curation and evaluation. Designed to support scalable experimentation across high- and low-resource languages with minimal human intervention.
-
A comparative evaluation against a production rule-based baseline. GPT-4o-mini and GPT-4o are benchmarked against an Apple Siri production rule-based normalization system, with overall WER and BLEU reported per language and per-category WER breakdowns in the appendix.
Main Findings
-
GPT-4o outperformed the production baseline on every language. Overall WER for GPT-4o ranged from 4.17% (German) to 7.88% (Japanese), versus baseline scores ranging from 9.72% (French) to 17.49% (Japanese). The paper notes the baseline predates the benchmark and was not tuned to it.
-
Both LLM systems improved over the baseline across the board. GPT-4o-mini's WER ranged from 6.60% (American English) to 11.19% (Mexican Spanish). Notably, GPT-4o-mini was slightly worse than the baseline on Lithuanian (10.54% vs. 10.04%), while GPT-4o reached 6.99% on that language.
-
BLEU gains were large. Baseline BLEU ranged from 47.76% (Japanese) to 70.84% (American English). GPT-4o reached 77.89% on Japanese and 89.85% on American English.
-
Per-language results (WER / BLEU). German: baseline 10.74 / 60.83, GPT-4o-mini 7.18 / 78.35, GPT-4o 4.17 / 84.92. American English: 9.84 / 70.84, 6.60 / 85.02, 4.28 / 89.85. Mexican Spanish: 11.92 / 55.74, 11.19 / 62.49, 7.69 / 71.90. French: 9.72 / 69.02, 7.70 / 79.84, 5.65 / 86.18. Italian: 15.02 / 55.14, 8.33 / 72.81, 4.56 / 84.96. Lithuanian: 10.04 / 67.12, 10.54 / 64.36, 6.99 / 73.96. Mandarin Chinese: 11.36 / 69.63, 6.65 / 79.04, 5.05 / 83.11. Japanese: 17.49 / 47.76, 10.78 / 68.32, 7.88 / 77.89.
-
Iterative prompt refinement measurably improved results. Appendix D reports GPT-4o WER by category for two prompt iterations. Between Iteration 2 and Iteration 3, overall WER dropped in every language — for example, Lithuanian from 12.22% to 6.99%, Japanese from 12.32% to 7.88%, and Mexican Spanish from 8.45% to 7.69%.
-
Context disambiguation is a key LLM advantage. The paper highlights cases a rule-based system needs explicit, verbose rules to handle: "3-2" is read as "three to two" in sports scores but digits are read individually in phone numbers, and a URL like
https://www.mediacityuk.co.ukis intelligently segmented into spoken components. -
Baseline error patterns were documented. Appendix B shows representative failures, including German address "Hauptstraße 45" normalized as "fünf und vierzig" instead of "fünfundvierzig", and French sports score "11-9" read as "onze moins neuf" instead of "onze à neuf".
-
Evaluation excluded non-meaning-affecting differences. Casing, optional spacing (such as in German compound words), and orthographic variants like "ß" versus "ss" were not counted as errors, so scoring focuses on differences that affect meaning, pronunciation, or intelligibility.
Methodology in Plain English
The researchers start from a Kestrel dataset subset — 14 categories drawn from the first file, which contains roughly 880,000 lines — and build a balanced English test set of 1,400 cases by keeping the first 100 examples per category and removing sentences without normalization targets. They initially tried translating the English data into target languages with NLLB-200, but found the translations unsuitable as ground truth due to inconsistent quality and inadequate localization of normalized forms.
Instead, they use DeepSeek-R1 to generate paired unnormalized and normalized text across 27 categories, with 20 examples per category (540 per language). Every example is then edited and verified by internal language experts, with orthography and formatting following target-language conventions — for instance, French uses a comma as the decimal separator with the currency symbol after the number.
For the normalization system itself, they use few-shot prompting rather than fine-tuning. Each prompt has three parts: an instruction prompt defining the task and listing the categories and rules, in-context learning examples showing how specific inputs should be normalized, and the target unnormalized input. The instruction prompt is standardized in English, while the ICL examples are localized per language. An 80-100 example ICL set — drawn from machine-translated Kestrel data and DeepSeek synthetic examples, all expert-validated — is kept separate from the benchmark. Japanese gets a supplementary prompt to guide katakana output.
To improve results, the team ran a hillclimbing loop: they compared LLM outputs against expert-verified development sets, identified recurring error patterns such as incorrect numeral expansions and inconsistent date or currency handling, then revised or added ICL examples targeting those failures. They also collaborated with language experts to flag inaccuracies, which fed back into benchmark corrections and prompt refinement.
Why This Matters
Impact on research. The paper shifts text normalization away from handcrafted rule authoring toward prompt engineering and example curation. It provides a standardized multilingual benchmark with 27 categories across eight languages, giving the field a shared evaluation target that earlier English- and GPT-focused work lacked. It also extends prior work by Zhang et al. (2024), which reported error rates roughly 40% lower than production WFST systems using GPT models in English, into a broader multilingual setting.
Real-world applications.
- Voice assistants and conversational agents, where user-facing text must be spoken correctly in real time.
- Screen readers and accessibility tools, where mispronounced dates, currency, and phone numbers degrade comprehension.
- Language learning applications, where incorrect normalization can teach learners wrong usage.
- Low-resource language TTS, where rule-based development is prohibitively expensive and where few-shot approaches lower the entry barrier.
Industry relevance. The work is authored at Apple and benchmarked directly against a production Siri rule-based system, indicating a deployment-oriented framing. The paper argues that rule-based pipelines demand ongoing annotation and patching, often taking months per language, while crowdsourcing normalization takes weeks of review cycles and can be expensive. Replacing that loop with prompt tuning and ICL is presented as a way to cut development overhead, accelerate iteration, and reduce reliance on large volumes of labeled data — relevant to any organization maintaining multilingual speech products.
Future Directions
-
Diacritization restoration for languages such as Arabic and Hebrew, where missing diacritics affect grammar and pronunciation, potentially addressed through targeted prompting.
-
Better handling of non-whitespace languages. The paper suggests integrated tokenizers and multitask learning for Japanese and similar languages, noting that normalizing to katakana resolves homograph ambiguity (e.g., 〇 read as "rei", "zero", or "maru") but may disrupt pitch accent patterns.
-
Incorporating suprasegmental features such as pitch accent and tone prediction into the normalization pipeline to preserve naturalness.
-
Broadening category coverage and robustness. The limitations section notes that ICL example categories may not capture the full range of normalization needs across languages and varieties, and that the model may still struggle with edge cases or uncommon linguistic patterns.
Target Audience
This paper is most useful to TTS and speech engineering practitioners who own normalization pipelines in production, to NLP researchers working on multilingual text processing and prompting strategies, and to teams building speech systems for lower-resource languages. It also suits applied researchers interested in how far in-context learning can substitute for rule-based engineering, though readers looking for model-architecture detail or training-time analysis will not find it here.
Authors’ abstract
Text Normalization (TN) is a key preprocessing step in Text-to-Speech (TTS) systems, converting written forms into their canonical spoken equivalents. Traditional TN systems can exhibit high accuracy, but involve substantial engineering effort, are difficult to scale, and pose challenges to language coverage, particularly in low-resource settings. We propose PolyNorm, a prompt-based approach to TN using Large Language Models (LLMs), aiming to reduce the reliance on manually crafted rules and enable broader linguistic applicability with minimal human intervention. Additionally, we present a language-agnostic pipeline for automatic data curation and evaluation, designed to facilitate scalable experimentation across diverse languages. Experiments across eight languages show consistent reductions in the word error rate (WER) compared to a production-grade-based system. To support further research, we release PolyNorm-Benchmark, a multilingual data set covering a diverse range of text normalization phenomena.