Research
Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
Overview Research area: Natural Language Processing — inductive biases of neural language models, artificial language (AL) learning, and syntactic typology. Technical level: Intermediate. Readers need

- arXiv
- 2510.12722
- Published
- 2025-10-14
- Authors
- Nadine El-Naggar, Tatsuki Kuribayashi, Ted Briscoe
AI summary
Overview
Research area: Natural Language Processing — inductive biases of neural language models, artificial language (AL) learning, and syntactic typology.
Technical level: Intermediate. Readers need basic familiarity with language model training/evaluation (perplexity, grammaticality judgments) and with syntactic formalisms; the paper's grammar machinery (categorial grammar, composition, coordination, permutation) is explained in the paper itself.
Scope: The paper builds 96 artificial languages with a Generalized Categorial Grammar (GCG) framework and tests whether simple RNN, LSTM, and Transformer models generalize better from short training sentences to longer unseen sentences in typologically frequent versus rare word orders.
What This Paper Is About
Earlier work (White and Cotterell 2021; Kuribayashi et al. 2024) used artificial languages generated by context-free grammars to ask whether language models find typologically common word orders easier to learn than rare or unattested ones. Those grammars were too limited to cover constructions such as object relative clauses, unbounded dependencies, and mildly context-sensitive (indexed language) structures like cross-serial dependencies, and they measured performance mainly on held-out data of the same length as training data. This paper replaces the grammar formalism with GCG, adds an extra word order parameter (covering VSO and OSV), and shifts the evaluation toward generalization to unseen longer sentences and grammaticality judgments rather than same-length perplexity.
Key Contributions
-
A GCG-based artificial language framework. The authors extend the context-free/PCFG AL formalization to Generalized Categorial Grammar (Wood 2014), using functional application, composition, coordination, and permutation (Briscoe 1997, 2000; they avoid CCG-style type raising for computational tractability). This allows the ALs to include object relative clauses and unbounded filler-gap dependencies.
-
A seventh word order parameter and 96 artificial languages. All parameters except
Ofollow White and Cotterell (2021); the newOparameter controls subject–object order and enables VSO and OSV word orders, yielding 96 distinct ALs. -
A length-generalization evaluation protocol. Training uses only a Short set (3–8 words), while testing uses held-out Short (3–8), Medium (9–10), and Long (11–20) sets, with additional targeted test sets (Recursive and Embedded relative clauses) and a minimal-pair grammaticality judgment task.
-
A comparison of three architectures under typological alignment metrics. Simple RNN, LSTM, and Transformer models are scored by the correlation between their performance and the real-world frequency of the corresponding word order (using WALS and, for complementizer statistics, Grambank).
Main Findings
-
Longer test sets make inductive bias clearer. In-domain (Short) perplexity distributions are comparatively flat, particularly for LSTMs, replicating a pattern also reported by White and Cotterell (2021); the PPLs on the Medium and Long test sets have larger variance and are therefore more informative about which word order a given LM handles well.
-
Typological alignment emerges mainly out-of-domain. Out-of-domain typological alignment (TA) scores are consistently negative, while the in-domain evaluation does not show this. Negative TA here means higher-probability (more common) word orders are easier for the model, so plausibility helps specifically with productive generalization rather than with fitting the training distribution. The authors note this in-domain result appears to contradict Kuribayashi et al. (2024), which did not consider length generalization, and suggest the difference may stem from this study's much shorter training sentences.
-
Architecture matters. The TA scores of the LSTM and RNN improved markedly in the out-of-domain evaluation, and the RNN on the Long test set achieved the best (lowest) correlation in all settings. The Transformer showed a good correlation only in-domain, which diminished out-of-domain. The authors connect this to working memory: the RNN has the most limited memory of the three (no LSTM gating, no attention-based context access), consistent with the idea that working-memory limits shape typologically frequent word orders.
-
Targeted complex constructions give phenomenon-dependent results. On the Recursive relative clause set, correlations were not statistically significant for any model (Transformer −5.1; LSTM 9.2; RNN 12.9), suggesting models may simply fail to learn such complex structures. On the Embedded relative clause set, correlations were negative and statistically significant for the Transformer (−23.5) and RNN (−18.1), with the LSTM at −3.7, broadly matching the earlier finding that typologically common ALs are easier to generalize.
-
Grammaticality judgments point the same way. All correlations between accuracy and typological plausibility were positive, but only the RNN was statistically significant in both settings (case type correlation 0.21; verb type correlation 0.23). Average accuracy was high in all models for case type (Transformer 97.7 ± 1.5, LSTM 97.2 ± 1.4, RNN 97.4 ± 1.4) but lower and more variable for verb type (Transformer 81.0 ± 14.7, LSTM 85.1 ± 9.6, RNN 77.4 ± 15.5).
-
Overall takeaway. Across experiments, recurrent models — especially the RNN — align best with typological distributions when generalization to longer sentences is required, while the Transformer is the least aligned in that setting.
Methodology in Plain English
The researchers generate synthetic languages rather than studying natural ones, so that a single grammatical factor (word order) can be varied in isolation while everything else is held fixed.
-
Define the grammar. Words are assigned GCG syntactic categories (11 categories total: NP, subject marker, object marker, adjective, transitive verb, intransitive verb, verb with complement, complementizer, preposition, relativizer, conjunction). Seven binary parameters decide the order of constituents:
S,VP,O,COMP,PP,ADJ,REL. For example,0101101corresponds to S=0, VP=1, O=0, COMP=1, PP=1, ADJ=0, REL=1 and gives basic English word order. -
Build grammatical templates. For each of the 96 parameter combinations, all sequences of word categories up to length 10 are parsed with a modified GCG chart parser (adapted from the NLTK CCGChartParser; type raising disabled, permutation added). A sequence counts as grammatical if it yields at least one derivation rooted in S.
-
Fill in words. Sentences are created by randomly sampling lexicons for each category, keeping length distributions uniform (e.g., 1K of length-3 sentences, 1K of length-4 sentences, …, 1K of length-8 sentences).
-
Create long sentences. Long templates (11–20 words) are made from short templates in three ways: concatenation, mid-sentence insertion with a conjunction, and appending with a conjunction. Valid results are filtered through the parser; 20,000 unique valid templates are sampled per AL, one sentence each.
-
Train and test. Models are simple RNN, LSTM, and Transformer, trained with Fairseq. Training data is 80K sentences of 3–8 words with uniform length distribution, early stopping with patience of five epochs and a maximum of 10,000 update steps. Three seeds are used and scores averaged.
-
Score. The central metric is typological alignment: Pearson correlation between model performance (perplexity, or grammaticality accuracy) and the real-world percentage of languages using that word order, based on WALS (with Grambank supplying complementizer statistics). Lower (more negative) TA in the perplexity case means more common word orders are easier for the model.
Why This Matters
The paper contributes to a live debate about whether neural language models have human-like inductive biases — and, by implication, whether they can be used as models of human language learning. It also shows that how you evaluate matters: measuring perplexity on same-length held-out data hides the very biases that show up when you demand generalization to longer sentences.
Real-world applications (implied by the framework and results):
- Evaluation design for LMs. It argues for out-of-domain, length-extrapolation test sets and targeted grammaticality probes as complements to standard perplexity, since holistic same-distribution scores can mask inductive bias.
- Architecture selection for morphologically or syntactically diverse languages. The finding that recurrent architectures align best with typologically common word orders on length generalization is relevant when choosing models for languages whose word order is under-represented in training data.
- Training data curation. The uniform-length sampling and template-coverage approach offers a way to reason about which sentence lengths a model actually learns to extend from.
- Cognitive modeling. The RNN's stronger alignment supports working-memory-limit accounts of why typologically frequent word orders exist, which is directly relevant to research using neural networks as hypotheses about human language.
Industry relevance: The core practical message is that architecture and evaluation protocol jointly determine whether a model handles longer inputs in unfamiliar orders. Teams deploying models on low-resource languages, on long documents, or on structured/controlled input (such as synthetic instruction formats) have a direct stake in whether a model's inductive bias favors the structures that actually occur in the target language.
Future Directions
-
Broaden the linguistic coverage. The authors state that their ALs do not differentiate verb tenses, do not include subject-verb agreement, do not model ambiguity, and assign each lexicon word to exactly one category; they call for systematic investigation of a broader range of linguistic phenomena within the framework.
-
Resolve the phenomenon-dependence in targeted evaluations. Recursive relative clauses produced no significant typological alignment while embedded relative clauses did; the authors note this requires further investigation with broader-coverage targeted evaluations.
-
Reconcile with prior contradictory results. The in-domain TA result appears to conflict with Kuribayashi et al. (2024); the authors attribute this tentatively to differences in training-data length distribution, which leaves open the question of how training sentence length interacts with measured inductive bias.
-
Extend and compare AL frameworks. The paper notes concurrent formulations such as dependency-based corpus modification (Xu et al. 2025), constituency-based non-adjacency (Hunter 2025), and multiple natural languages as seeds (Yang et al. 2025), inviting comparison across these approaches.
Target Audience
Researchers in computational linguistics and NLP who work on language model inductive bias, artificial language learning, and syntactic generalization; cognitive scientists interested in whether neural networks can account for typological word order distributions; and machine learning practitioners evaluating model behavior on inputs longer than those seen in training. Readers most likely to benefit are those already comfortable with perplexity-based evaluation and basic syntactic formalisms, since the paper's novelty lies in the experimental setup rather than in new modeling techniques.
Authors’ abstract
Whether language models (LMs) have inductive biases that favor typologically frequent grammatical properties over rare, implausible ones has been investigated, typically using artificial languages (ALs) (White and Cotterell, 2021; Kuribayashi et al., 2024). In this paper, we extend these works from two perspectives. First, we extend their context-free AL formalization by adopting Generalized Categorial Grammar (GCG) (Wood, 2014), which allows ALs to cover attested but previously overlooked constructions, such as unbounded dependency and mildly context-sensitive structures. Second, our evaluation focuses more on the generalization ability of LMs to process unseen longer test sentences. Thus, our ALs better capture features of natural languages and our experimental paradigm leads to clearer conclusions -- typologically plausible word orders tend to be easier for LMs to productively generalize.