Research
Function Words as Statistical Cues for Language Learning
Overview Research area: Natural Language Processing, specifically computational language acquisition, corpus linguistics, and the interpretability of transformer language models. Technical level: Inte

- arXiv
- 2601.21191
- Published
- 2026-01-29
- Authors
- Xiulin Yang, Heidi Getz, Ethan Gotlieb Wilcox
AI summary
Overview
- Research area: Natural Language Processing, specifically computational language acquisition, corpus linguistics, and the interpretability of transformer language models.
- Technical level: Intermediate. The paper combines Universal Dependencies corpus statistics with controlled transformer training and probing, so familiarity with dependency parsing, language modeling, and attention is helpful but not essential.
- One-sentence scope: A cross-linguistic corpus study of 186 languages plus counterfactual language-model training experiments on English testing whether three distributional properties of function words (high frequency, reliable structural association, phrase-boundary alignment) support syntactic learning.
What This Paper Is About
Researchers have long argued that function words (determiners, auxiliaries, prepositions) help learners extract grammar from linear speech because they are frequent, reliably tied to specific syntactic structures, and consistently placed at phrase boundaries. Those claims, however, rested mostly on English and on small artificial languages, leaving open whether the properties generalize across languages and how each one contributes to learning. This paper tests all three properties at scale, then systematically destroys each one in training data to see how much syntactic learning degrades.
Key Contributions
- A cross-linguistic corpus analysis using Universal Dependencies (v2.17, 339 treebanks across 186 languages) showing that all three distributional properties of function words hold universally in the sample.
- A set of controlled "counterfactual" language variants built from Wikipedia text in which each property is separately manipulated, used to train GPT-2 Small models from scratch and 5-gram baselines.
- Evidence for a Goldilocks effect: function words must be frequent enough to be reliable but diverse enough to remain structurally informative.
- Attention probing and function word masking/deletion ablations showing that different training conditions produce systematically different internal reliance on function words.
Main Findings
- All three properties are universal in the sample. Across 186 languages, function words occupy a small inventory but a disproportionately large share of tokens (points above the diagonal in Figure 1a), while content words fall below the diagonal. Function words also show lower syntactic dependency entropy than content words across all languages (Figure 1b).
- Function words align with phrase boundaries. The median ratio of function words at phrase boundaries is 0.95 (minimum 0.55 in Korean), compared to only 0.58 for content words.
- Intact function words yield the best syntactic generalization. On BLiMP, the NaturalFunction transformer condition reaches 72.7 overall accuracy, while NoFunction drops to 60.7 (-12.0) and every other manipulated condition falls in between the two.
- Disrupting any property hurts learning. A linear mixed-effects model (accuracy ~ condition + (1|category:phenomenon) + (1|seed)) shows all conditions have significant negative effects relative to NaturalFunction (p < 0.05), except FiveFunction, which is marginally worse (p = 0.08).
- Rare and collapsed inventories both hurt. FiveFunction (5 types) reaches 70.9 (-1.8) and MoreFunction (roughly 1.2k types) reaches 69.7 (-3.0), compared to 72.7 for NaturalFunction — the Goldilocks pattern.
- Structural association matters most. RandomDep (67.0, -5.7) and BigramDep (67.4, -5.3) produce larger drops than frequency manipulations (FiveFunction -1.8, MoreFunction -3.0) or the boundary manipulation WithinBoundary (69.7, -3.0).
- n-gram models trail transformers everywhere. The 5-gram baseline scores 55.5 on NaturalFunction versus 72.7 for the transformer, and it improves on BigramDep (56.1) over NaturalFunction (55.5), indicating transformers capture structure unavailable to linear heuristics.
- Category effects are uneven. Filler Gap tests improve notably in FiveFunction (72.5, +6.6 over NaturalFunction's 65.9), while S-V Agreement and Irregular forms take the largest hit in NoFunction (-14.5 and -30.6 respectively).
- NaturalFunction produces concentrated function-word attention heads. A small number of heads, primarily in Layers 3 and 4, account for most function word attention across BLiMP categories; RandomDep, BigramDep, and WithinBoundary do not produce similarly concentrated patterns. NaturalFunction also shows lower mean entropy and standard deviation across three random seeds.
- Deletion hurts more than masking. Removing function words entirely causes a larger accuracy drop than blocking attention to and from them. Under deletion, NaturalFunction shows the largest drop, followed by FiveFunction and WithinBoundary. Under masking, MoreFunction shows the largest drop.
- Disrupted association reduces reliance. BigramDep and RandomDep show the smallest performance drops in both ablation experiments. Deletion drops are significant in all conditions (p < 0.01); masking drops are significant in most conditions (p < 0.05) except BigramDep and MoreFunction, and masking results are less stable across runs.
Methodology in Plain English
The authors first check whether claims about function words hold beyond English. Using Universal Dependencies parses, they measure three things per language: how many tokens relative to types function words account for, how predictably a function word's syntactic neighbors cluster (dependency entropy), and how often function words sit at the left or right edge of a subtree yield (an approximation of a phrase boundary, with adjustments for nested structures and for categories like pronouns, particles, and auxiliaries that behave unreliably in dependency-to-constituent mapping).
Second, they build controlled versions of English Wikipedia text. Content words and total dataset size are held constant, and sentence-length distributions are matched across conditions except NoFunction. All text is lowercased and parsed with Stanza. Function words are defined from closed-class tags (det, adp, cconj, sconj, aux), and the inventory of 116 English function words comes from the GUM and EWT treebanks after filtering items with fewer than 10 occurrences. Lexical frequency is varied by keeping zero function words (NoFunction), collapsing each category to one type (FiveFunction, 5 types total), using the natural inventory (Standard Function / NaturalFunction, 116 types), or expanding each item into 10 pseudowords with Wuggy (MoreFunction, about 1.2k types). Structural predictability is varied by making function word identity depend on the following word (BigramDep) or by randomly shuffling identities while preserving location (RandomDep), against the natural PhraseDependency baseline. Boundary alignment is varied by moving function words next to their syntactic heads (WithinBoundary) versus keeping them at boundaries (AtBoundary); this changes the location of function words in 55% of positions across over 99% of training sentences.
Third, they train GPT-2 Small models from scratch on each variant with a dedicated tokenizer, for 10 epochs, averaged over 3 random seeds, plus 5-gram KenLM baselines with Kneser-Ney smoothing. They avoid perplexity because the manipulations change corpus entropy, and instead evaluate on BLiMP, applying the same transformations to the test sets and filtering out categories whose critical word is a function word, pairs that become identical, and using intersection filtering so all models see the same phenomena.
Finally, they probe attention heads using the method of Aoyama and Wilcox (2025) to find heads that consistently attend to function words, and they run two ablations: masking attention to and from function word tokens at evaluation time, and deleting function words entirely by evaluating on NoFunction BLiMP.
Why This Matters
The paper moves claims about function words from English and artificial languages onto a 186-language empirical footing and shows that not all statistical cues contribute equally: reliable structural association exerts a stronger effect than boundary alignment or frequency. It also gives a mechanistic account, linking training-data properties to the attention heads that emerge and to the degree to which models depend on function word information at test time.
Real-world applications:
- Language education and curriculum design. Understanding that frequent, structurally reliable function words anchor grammar learning could inform how early reading and second-language materials sequence vocabulary.
- Low-resource language technology. If these properties are universal, tools that bootstrap structure from function words may transfer to languages with limited parsed data.
- Tokenizer and pretraining decisions for NLP systems. The Goldilocks finding implies that aggressively shrinking or massively expanding the function word vocabulary can both degrade downstream syntactic ability.
- Speech and language interfaces. Since function words are acoustically reduced and the paper's experiments are text-only, systems that must recover them from noisy audio face an open question about whether the same cues survive.
Industry relevance: teams training or fine-tuning language models, building multilingual pipelines, or developing educational and clinical language-assessment tools have a concrete, testable hypothesis about which vocabulary distributions support grammar acquisition and which internal representations to expect.
Future Directions
- Extending the counterfactual modeling experiments beyond English to typologically diverse languages, including whether effect magnitudes vary with word order or morphological richness.
- Investigating grammatical morphology, since languages like Turkish encode function-like information in bound morphemes rather than free function words, and Universal Dependencies does not annotate the morphology level.
- Testing whether the same patterns hold for speech input, where prosodic cues such as stress and rhythm may supply additional structural information that the written-text experiments omit.
- Clarifying the status of the MoreFunction condition, whose pseudowords lack natural collocational histories, and developing evaluation beyond BLiMP as a proxy for syntactic generalization since a 5-gram model can score above chance.
Target Audience
Researchers in computational linguistics and language acquisition, cognitive scientists studying statistical learning, and NLP practitioners interested in how pretraining data composition shapes syntactic generalization. The paper is most useful to readers comfortable with dependency syntax and transformer training, though the cross-linguistic findings and the Goldilocks result are accessible to a broader audience.
Authors’ abstract
What statistical properties might support learning abstract grammatical knowledge from linear input? We address this question by examining the statistical distribution of function words. Function words have been argued to aid acquisition through three distributional properties: high frequency, reliable syntactic association, and phrase-boundary alignment. We conduct a cross-linguistic corpus analysis of 186 languages, which confirms that all three properties are universal. Using counterfactual language modeling and ablation experiments on English, we show that preserving these properties facilitates acquisition in neural learners, with a Goldilocks effect: function words must be frequent enough to be reliable, yet diverse enough to remain informative to structural dependency. Probing analyses further reveal that different learning conditions produce systematically different reliance on function words.