Research
The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
Overview Research area: Natural Language Processing / interpretability of syntax in Transformer-based language models (TLMs). Technical level: Intermediate. The paper is a systematic review rather tha

- arXiv
- 2601.19926
- Published
- 2026-01-09
- Authors
- Nora Graichen, Iria de-Dios-Flores, Gemma Boleda
AI summary
Overview
Research area: Natural Language Processing / interpretability of syntax in Transformer-based language models (TLMs).
Technical level: Intermediate. The paper is a systematic review rather than a new model or method, but it assumes familiarity with concepts such as probing classifiers, minimal-pair benchmarks, and mechanistic interpretability.
Scope in one sentence: A systematic, quantitative synthesis of 337 articles and 3,074 annotated datapoints covering what is known about syntactic knowledge in Transformer-based language models, how that research has been conducted, and where it remains thin.
What This Paper Is About
Since at least Gulordava et al. (2018), researchers have noticed that language models trained only to predict tokens seem to pick up a substantial amount of syntax, but the field had no general picture of what exactly is learned or how it is used. The authors perform a systematic review — a format with transparent search, selection, and annotation criteria that yields quantitative evidence — of the literature on syntactic knowledge in Transformer-based language models. The goals are twofold: map the research landscape (what has been done and how) and summarize what is currently known about syntactic knowledge in TLMs.
Key Contributions
-
An annotated database of 337 articles. The reviewed studies are compiled into a publicly available annotated database with 32 annotation variables across four categories (meta-data, model-related information, experimental design, interpretability techniques), released together with the analysis code at https://github.com/norgrai/syntax_review. The database is designed to remain open to future extensions via a dedicated submission form.
-
A quantitative, systematic synthesis rather than a narrative review. The paper contrasts itself with six previous narrative reviews (Limisiewicz and Mareček, 2020; Linzen and Baroni, 2021; Kulmizev and Nivre, 2022; Chang and Bergen, 2023; Millière, 2024; López-Otal et al., 2025) and provides a transparent mapping between available evidence and conclusions across 3,074 datapoints, over 3,000 individual results.
-
A bottom-up typology of syntactic phenomena. Starting from the terminology used in the studies themselves, the authors group work into 11 syntactic categories recognizable to both linguistics and NLP communities; the typology is offered as a standardization resource for the field.
-
Concrete recommendations for future research. The paper recommends fuller data reporting (including architecture, size, tokenization, training data, task framing, and prompting), greater standardization of how syntactic phenomena are defined, multilingual benchmarks that go beyond translations or adaptations of English datasets, and more causal/interventional methods combined with observational ones.
Main Findings
-
TLMs encode a non-trivial amount of syntactic knowledge. Behavioral, probing, and mechanistic evidence collectively support this. TLMs average 72% on the English BLiMP benchmark against a 50% random baseline, and between 79% and 85% of the analyzed studies identify syntactic knowledge in TLMs.
-
Formal syntax is easier than the syntax-semantics interface. All agreement categories in BLiMP have medians over 85%, while binding, argument structure, NPI licensing, control/raising, quantifiers, and island effects all have medians below 75%. A mixed-effects regression on non-benchmark scores confirms this pattern: POS and other lexical properties, and agreement and feature-checking, have higher-than-average accuracies, while hierarchical structure, negation and NPIs, and argument structure and constructions have lower accuracies.
-
One exception: ellipsis. Several TLMs match or outperform humans on the ellipsis category of BLiMP, despite ellipsis depending on syntax-semantics interactions. The authors attribute this to dataset characteristics — BLiMP evaluates only specific cases of noun phrase ellipsis whose violations may be comparatively easy to detect — rather than to ellipsis being easy for models.
-
Less digitally supported languages perform worse. Language support is a significant predictor of accuracy, with each higher support category (using the 5-level classification of Joshi et al. (2020)) associated with a gain of approximately 7 percentage points (β = 0.065, p < 0.001).
-
Syntactic ability looks general, not phenomenon-specific. Model rankings across phenomena are highly consistent; the large within-phenomenon variance is mostly driven by how good each model is at syntax overall, as measured by its overall BLiMP score.
-
Scaling helps. Larger TLMs and TLMs trained on more data generally achieve higher performance, and bidirectional models tend to outperform causal models on average BLiMP scores.
-
Syntactic information is concentrated in middle layers. Of the 56 probing-based studies reporting localization information, middle layers are most commonly identified as showing the strongest probing performance. Studies typically describe these layers as showing stronger, not exclusive, encoding.
-
Probing success does not track behavioral difficulty. 85% of the 147 probing-based studies examined report successfully recovering syntactic information from model representations, and no substantial differences appear across syntactic phenomena — which contrasts with behavioral results and supports the view that successful probing alone is not a reliable measure of the quality or functional use of representations.
-
Mechanistic evidence is positive but narrow. 79% of the 140 mechanistic papers in the database find clear supporting evidence for syntactic knowledge, but the vast majority of that evidence comes from studies focusing exclusively on English; syntax-semantics interface phenomena have low coverage in the mechanistic literature.
-
The field over-focuses on English and a few models. 69% of articles are on English only and 91% include English. BERT alone accounts for 30% of results; BERT plus its variants account for 58%; BERT, GPT-2, RoBERTa, and mBERT together account for 62%. 66% of studies evaluate monolingual models (the vast majority English), 20% multilingual models, and 14% both.
-
Methods are diverse but mostly observational. Behavioral, probing, and mechanistic methods are relatively balanced across the years, but only 19% of studies use interventional methods, and 81% rely on purely observational methods. Probing has declined substantially — for instance, probing studies with bidirectional TLMs dropped from 27 in 2022 to just 11 in 2023.
-
Benchmark coverage is thin. Only 42 studies used benchmarks, 29 of which used BLiMP. The others were the CoLA task from GLUE (11 uses), SyntaxGym (3 uses), and Holmes and HANS (1 use each). Only 11 of the 25 BLiMP papers report per-category results.
-
Mechanistic work is diverse. The authors find over 80 specific methods in mechanistic studies, which they view positively for a young field. Post-hoc analysis of fully trained models at a global level clearly dominates, though a healthy number of papers track development and check both global and local trends.
-
Publication trends. Interest took off after BERT appeared, peaked in 2022, and decreased afterward (with the caveat that 2025 data covers only through July). Initially most work focused on bidirectional TLMs; causal models later dominated.
Methodology in Plain English
The authors followed a systematic review protocol in three stages: identifying candidate papers, assessing eligibility, and annotating the survivors. Data collection took place in summer 2025 with a cut-off date of July 31, 2025.
To find candidates, they used two strategies: a keyword search in Google Scholar combining terms like "syntactic structure/knowledge" with "LLMs/LMs/transformer," and snowballing from the six prior narrative reviews. The snowballing screened all 263 references cited by those reviews, some of the articles cited within those 263, and most articles citing the reviews themselves. This produced 636 candidate articles.
Eligibility screening excluded studies on non-Transformer or encoder-decoder architectures, studies that analyze TLMs without empirically assessing syntactic knowledge, publications with no clear empirical component (position papers and surveys, including the prior reviews), and non-peer-reviewed, non-conference, non-archived-preprint publications such as blog posts and theses. This left 337 articles.
Annotation used 32 variables in four categories. Twenty-eight variables were annotated manually; four (model type, syntactic phenomena, experimental materials, method name) used a validated semi-automated AI-assisted workflow, with final labels determined manually. For the analysis, the authors classify methods as behavioral, probing, or mechanistic following Millière (2024), fit a mixed-effects regression with syntactic phenomena, model type (causal/bidirectional), and language support as fixed effects and model name and study ID as random effects, and code studies for author-reported strength of evidence.
Why This Matters
The review gives the field its first quantitative map of a large and methodologically fragmented literature, letting researchers see which conclusions are well supported, which rest on a narrow English and BERT-centric base, and where evidence is missing. It also provides a shared typology and an open, expandable database, which are prerequisites for meaningful cross-study comparison.
Real-world applications that follow from the findings (the paper itself does not enumerate applications):
- Multilingual product development. The finding that performance drops with lower digital support flags where NLP systems are likely to fail for under-resourced languages, informing where to invest data collection.
- Benchmark and evaluation design. The observation that BLiMP focuses on cases expressible as minimal pairs and aggregates heterogeneous cases into equally weighted composite scores is directly relevant to anyone designing evaluation suites or interpreting leaderboard numbers.
- Trust and failure analysis. Knowing that models rely on surface-level heuristics and are weaker at the syntax-semantics interface — binding, negation, ellipsis, filler-gap dependencies — helps anticipate error modes in downstream text processing.
- Research infrastructure and reproducibility. The call for complete data reporting (architecture, size, tokenization, training data, task framing, prompting) and for standardized multilingual benchmarks addresses a practical bottleneck for anyone trying to compare models.
Industry relevance: Model selection for syntactic tasks currently rests on a literature skewed toward English and a handful of architectures, so conclusions about robustness and generality are limited. The review also argues that computing cost is being spent on observational analysis while causal methods that could explain why representations emerge remain scarce — a gap relevant to anyone trying to intervene on model behavior rather than merely measure it.
Future Directions
-
Expand coverage of the syntax-semantics interface. Phenomena such as binding, coreference, negation, ellipsis, and filler-gap dependencies have received very little attention, even though models generally perform worse on them than on more formal aspects of syntax.
-
Build standardized multilingual benchmarks. The authors argue for benchmarks that go beyond direct translations or adaptations of English datasets, citing the Spanish variant of SyntaxGym (Pérez-Mayos et al., 2021c) — which incorporates language-specific features like subjunctive mood and flexible word order alongside English-compatible phenomena — as a "best of both worlds" model.
-
Increase causal and interventional work. Only 19% of studies use interventional methods. The authors encourage interventional approaches while emphasizing that no single method suffices and that the strongest evidence comes from combining observational and interventional paradigms and comparing different architectures.
-
Study how syntax emerges during training. The field has mostly analyzed trained models at a global level; more attention is needed on the emergence of syntactic knowledge during training and on the connections between local and global representations and mechanisms.
-
Open theoretical questions. Whether TLMs implement near-symbolic mechanisms (as hypothesized by Boleda, 2025); when and how TLMs fall back on superficial statistical patterns, since performance on agreement, binding, or scope often degrades for rare or nonce lexical items; and whether improvements with model scale reflect genuine generalization or massive memorization of specific instances (Lake and Baroni, 2018).
Target Audience
Researchers and graduate students in NLP interpretability, computational linguistics, and psycholinguistics who want a quantitative map of the syntax-in-TLMs literature and a reusable annotated database. It is also useful for practitioners who need to interpret benchmark results critically, and for methodologists and reviewers interested in how systematic reviews can be conducted in computational linguistics — a format the authors note is still uncommon in the field. The paper's Limitations section notes that the review focuses on Transformer-based models and textual input, so findings may not generalize to recurrent or other neural architectures, and that the keyword search was conducted in English, likely under-retrieving publications written in other languages.
Authors’ abstract
We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on over 3,000 datapoints spanning a wide range of syntactic phenomena, languages, models, and methods. We take the data to collectively show that TLMs encode a non-trivial amount of syntactic knowledge. Behavioral evidence shows strong performance on formal syntactic phenomena, but weaker and more variable performance on phenomena at the syntax-semantics interface. Performance is also consistently lower for languages with less digital support. Probing and mechanistic studies further support the presence of syntactic knowledge in TLMs. Yet, because most work remains observational and methodologically heterogeneous, insight into the detailed computational mechanisms underlying syntactic processing remains limited. At the same time, the literature remains heavily concentrated on English and BERT-like models. We discuss the implications of our results and provide recommendations for future research.