Skip to content
AI.info

Research

Corpus Frequencies in Morphological Inflection: Do They Matter?

Overview Research area: Natural Language Processing, specifically morphological inflection (generating inflected word forms from a lemma and a set of grammatical tags). Technical level: Intermediate.

Corpus Frequencies in Morphological Inflection: Do They Matter?
arXiv
2510.23131
Published
2025-10-27
Authors
Tomáš Sourada, Jana Straková

AI summary

Overview

Research area: Natural Language Processing, specifically morphological inflection (generating inflected word forms from a lemma and a set of grammatical tags).

Technical level: Intermediate. The paper assumes familiarity with sequence-to-sequence models and train/dev/test methodology, but its central idea — weighting training examples by how often they occur in a corpus — is conceptually simple.

Scope: The paper investigates whether corpus frequency information should be incorporated into three stages of a morphological inflection pipeline — data splitting, evaluation, and training-time sampling — across 43 languages using Universal Dependencies corpora.

What This Paper Is About

Morphological inflection systems are traditionally trained and evaluated on lexicons of lemma-tag-form triples treated as equally important, with no regard for how often each form actually appears in real text. That is unrealistic for deployment: users type frequent words far more often than rare ones. The paper asks whether injecting corpus frequency information into the split, the evaluation metric, and the training sampling procedure improves a system's ability to predict the forms people actually use.

Key Contributions

  1. A combined split strategy. The authors use a train-dev-test split that is simultaneously lemma-disjoint (so the model's generalization to unseen lemmas can be measured) and frequency-weighted (so the distribution of items across frequency bands resembles real text).

  2. Token accuracy as a deployment-oriented metric. Alongside standard type accuracy, which weighs every item equally, the paper reports token accuracy, which weights items by corpus occurrence count and better approximates performance on running text. The authors note this metric was previously proposed by Nicolai et al. (2020) but never used in inflection work.

  3. Frequency-aware training. The paper introduces a corpus-frequency temperature τ, a real-valued hyperparameter that controls how strongly training examples are sampled in proportion to their corpus counts, spanning uniform sampling (τ = 0) to raw-frequency sampling (τ = 1) on a continuous scale.

  4. Open-source release. Code is publicly released at https://github.com/tomsouri/corpus-frequencies-in-inflection.

Main Findings

  • Frequency-aware training wins in most languages. The abstract states that frequency-aware training outperforms uniform sampling in 26 out of 43 languages.

  • A middle-ground temperature is best on average. In dev-set token accuracy (Table 2), the macro average improves continuously as τ rises to 0.5 and then declines, making τ = 0.5 — sample weights equal to the square root of raw occurrence counts — the best single value across all languages, at 86.02% macro-average token accuracy.

  • Extreme temperature causes collapse. τ = 2.0 leads to a complete failure in essentially every language, with a macro-average token accuracy of 17.54% and type accuracy of 12.86%. The sole exception is Old French, where τ = 2.0 yields the best performance.

  • Languages differ in their preferred temperature. Basque, Old French, and Turkish improve steadily up to τ = 1.0. Belarusian, Bulgarian, Croatian, Dutch, German, Gothic, Irish, Italian, Latin, Polish, Pomak, Portuguese, Russian, Spanish, and Welsh peak near τ = 0.5. Afrikaans, Catalan, Czech, Danish, Estonian, Finnish, French, Greek, Hungarian, Icelandic, Latvian, Manx, Norwegian, Romanian, Scottish Gaelic, Slovenian, and Swedish show little benefit, with best results near τ = 0. English, Galician, and Slovak do best with negative τ values, which emphasize rare forms. Breton, Lithuanian, Low Saxon, and Ukrainian show noisy, inconclusive patterns.

  • Type accuracy follows similar trends. Even though type accuracy ignores frequencies, the authors observe comparable temperature patterns in most languages, which they find surprising and discuss further in Section 5 of the paper.

  • Token and type accuracy rank models similarly. Despite absolute differences — for example English at τ = 0.0 reaches 96.22% token accuracy versus 93.37% type accuracy — the two metrics largely agree on model ordering under these splits.

  • Checkpoint-selection metric matters little. Training one system with checkpoints chosen by token accuracy and another by type accuracy produced negligible losses: the largest absolute drop was 0.9% in Breton, followed by 0.7% in Basque, 0.3% in Spanish and English, and 0.2% in Czech.

  • The tuned model beats prior baselines. The tuned Transformer without frequency-aware training outperforms both SIGMORPHON neural baselines (Wu et al., 2021) in all five development languages.

  • Copy baseline is far weaker on test. In the test comparison, the naive copy baseline that echoes the lemma scores 72.19% token accuracy on Afrikaans (AfriBooms), 45.23% on Basque (BDT), and 43.44% on Belarusian (HSE), versus roughly 85–89% for trained systems.

  • Test results. For Afrikaans, τ = 0.5 reaches 86.41% token accuracy and 82.93% type accuracy, τ_best reaches 86.31% token and 83.82% type, and τ = 0.0 reaches 85.89% token and 82.61% type. For Basque, τ_best reaches 89.59% token and 87.57% type.

Methodology in Plain English

The authors needed corpus frequencies, which the commonly used UniMorph lexicon does not provide. They considered three options: use an annotated corpus such as Universal Dependencies (UD) directly, align a raw corpus with UniMorph and discard uncovered entries, or align and assign a small constant count to uncovered entries. They chose the first option as the most straightforward, accepting lower lemma/form coverage, and extracted unique lemma-tag-form triples with occurrence counts from UD corpora for 43 languages. For hyperparameter tuning they used five development languages: Czech, English, Spanish, Breton, and Basque.

The canonical UD split overlaps on lemmas and items, so they resplit the data. Their procedure counts total lemma occurrences, samples lemmas into the training set weighted by occurrence count, samples remaining lemmas uniformly into the dev set, and places the rest in test, targeting an 8:1:1 ratio in total occurrence counts.

They trained a small-capacity encoder-decoder Transformer from scratch, taking the lemma-tag pair as input and predicting the inflected form. Hyperparameters included 3 layers, layer dimension 256, 4 attention heads, feed-forward dimension 64, batch size 512, learning rate 0.001, cosine decay, max 960 epochs, trained with Adam; checkpoints were selected on dev token accuracy per language.

For frequency-aware training, each training item receives a sample weight equal to its raw corpus count raised to the power τ, and batches are drawn by weighted random sampling. τ = 0 reproduces uniform sampling, τ = 1 gives weights equal to raw counts, and τ = 0.5 gives square-root weights — so an item occurring 400 times is sampled 20 times more often than one occurring once. At τ = 2 the same pair would differ by a factor of 160,000. The authors tested τ values across [0, 1] plus 1.1, 2, and negative values.

Why This Matters

Impact on research. The paper challenges a decades-old convention of treating inflection lexicons as frequency-blind. It connects methodological rigor (lemma-disjoint splits from Goldman et al., 2022 and 2023) with deployment realism (frequency weighting from Kodner et al., 2023 and Sourada and Straková, 2025), and it revives token accuracy, a metric proposed in 2020 but never previously applied to inflection.

Real-world applications:

  • Writing assistants and grammar checkers that need to inflect the words users actually type, which skew heavily toward frequent vocabulary.
  • Text generation and machine translation systems that must produce inflected forms of common words correctly in fluent output.
  • Speech recognition post-processing and dictation tools that reconstruct written forms from spoken input.
  • Language learning and documentation tools, where frequency-aware behavior may align with how learners encounter a language.

Industry relevance. Any deployed inflection component sees a frequency-skewed input distribution, so optimizing for token accuracy rather than type accuracy is directly relevant to product quality. The finding that weighting training data by roughly the square root of frequency helps across many languages — and the availability of open-source code — gives engineering teams a concrete, low-cost tuning knob.

Future Directions

  • Explore the alternative data-source strategies. The authors explicitly leave for future work the options of aligning a raw corpus with UniMorph and discarding zero-count entries, and of assigning small constant counts to theoretically possible forms not found in the corpus — both of which would raise lemma and form coverage compared to using UD alone.

  • Understand why type accuracy also benefits from frequency-aware training. The paper poses this as an open question, since the effect is unexpected when evaluation disregards frequencies.

  • Investigate the token-versus-type accuracy relationship further. The authors state they have not shown a substantial difference between the two metrics in terms of model ordering under their split and evaluation conditions, and call for more exploration.

  • Explain language-specific behavior. Why some languages prefer τ near 0, others τ = 1, and still others negative τ — and why Old French uniquely benefits from τ = 2.0 — remains unresolved.

Target Audience

Researchers and engineers working on morphological inflection, low-resource morphology, and morphologically rich languages, especially those familiar with the SIGMORPHON shared task tradition and Universal Dependencies. It is also relevant to NLP practitioners who deploy inflection components and need to decide how to weight training data and which evaluation metric reflects their production traffic. Readers interested in experimental methodology — dataset splitting, sampling, and metric choice — will find the systematic 43-language comparison useful.

Authors’ abstract

The traditional approach to morphological inflection (the task of modifying a base word (lemma) to express grammatical categories) has been, for decades, to consider lexical entries of lemma-tag-form triples uniformly, lacking any information about their frequency distribution. However, in production deployment, one might expect the user inputs to reflect a real-world distribution of frequencies in natural texts. With future deployment in mind, we explore the incorporation of corpus frequency information into the task of morphological inflection along three key dimensions during system development: (i) for train-dev-test split, we combine a lemma-disjoint approach, which evaluates the model's generalization capabilities, with a frequency-weighted strategy to better reflect the realistic distribution of items across different frequency bands in training and test sets; (ii) for evaluation, we complement the standard type accuracy (often referred to simply as accuracy), which treats all items equally regardless of frequency, with token accuracy, which assigns greater weight to frequent words and better approximates performance on running text; (iii) for training data sampling, we introduce a method novel in the context of inflection, frequency-aware training, which explicitly incorporates word frequency into the sampling process. We show that frequency-aware training outperforms uniform sampling in 26 out of 43 languages.

Read the original paper