Skip to content
AI.info

Research

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

Overview Research area: Psycholinguistics and natural language processing — specifically, how word predictability is quantified and how those estimates predict human reading behavior. Technical level:

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
arXiv
2601.09886
Published
2026-01-14
Authors
Sathvik Nair, Byung-Doh Oh

AI summary

Overview

Research area: Psycholinguistics and natural language processing — specifically, how word predictability is quantified and how those estimates predict human reading behavior.

Technical level: Intermediate. Readers should be comfortable with concepts like surprisal, cloze probability, linear mixed-effects regression, and language model probabilities. The paper is framed as a psycholinguistic methods paper rather than an NLP systems paper.

Scope: The paper re-establishes that GPT2 surprisal predicts reading times better than cloze surprisal, then runs a set of targeted manipulations on GPT2's probabilities to identify exactly why that advantage exists.

What This Paper Is About

Word predictability can be measured two ways: by asking humans to complete a sentence (the cloze task) or by reading probabilities off a language model. Recent work has found that LM surprisal (negative log probability) predicts reading times better than cloze surprisal, leading some to argue that LMs should replace the cloze task entirely. Nair and Oh argue that fit-to-data alone is not sufficient justification, since different predictability measures can support different scientific conclusions about prediction during comprehension. Their goal is to determine why LM surprisal wins, by surgically degrading GPT2's probabilities so they resemble cloze data in specific ways.

Key Contributions

  1. A systematic re-evaluation of cloze surprisal across four reading time datasets, testing six smoothing factors (V = 50, 100, 200, 500, 1000, 2000) crossed with six functional forms of cloze probability, yielding 36 regression models fitted to about 50% of observations for each of six measures.

  2. A confirmation that GPT2 surprisal outperforms the best-fitting cloze surprisal on four out of six reading time measures, and that GPT2 surprisal subsumes the effect of cloze surprisal (not the reverse).

  3. A novel intervention methodology: rather than comparing measures as-is, the authors manipulate GPT2's probabilities to test three distinct hypotheses about the source of its advantage — resolution (H1), semantics (H2), and frequency (H3).

  4. A negative result on methods for combining cloze responses with LM probabilities, via similarity-adjusted (SA) surprisal based on cloze responses and on GPT2 samples.

Main Findings

  • GPT2 surprisal is the stronger predictor, but not everywhere. On four of the six measures, GPT2 surprisal predicted reading times better than cloze surprisal while the reverse never held. The two self-paced reading (SPR) measures showed no significant difference and were not well predicted by either predictor, which the authors suggest may reflect task-based differences between SPR and eyetracking.

  • Cloze probability is not linear in reading time. Contrary to the linearity reported by Brothers and Kuperberg (2021), across the broader set of datasets the best-fitting configuration was the quadratic transform S(w_t)^2 with a smoothing factor of V = 200. Transforming probabilities into surprisal notably improved fit; the smoothing factor and the power transforms had smaller effects.

  • Hypothesis 1 (Resolution) is supported. When GPT2 probabilities were replaced by counts of sampled words, using the same number of samples N as cloze responses and the same add-one smoothing form, the resulting surprisal became a weaker predictor than unaltered GPT2 surprisal. Median performance over five runs was reported.

  • Hypothesis 2 (Semantics) is supported. Collapsing GPT2's vocabulary into k-means clusters and assigning each word its cluster's total probability (clustering performed with token embeddings, k in {20, 40, 80, 100, 500, 1000}, median over five runs) significantly reduced fit to reading times. Results are shown for 80 clusters as a representative setting.

  • Hypothesis 3 (Frequency) is supported. Zeroing out infrequent GPT2 subword tokens and renormalizing over frequent tokens (thresholds of 10^3, 10^4, and 10^5 occurrences per billion words from wordfreq; 10^4 reported as representative) significantly reduced fit.

  • The degraded variants lose to cloze on Provo. In the by-measure breakdown, the drop in fit was most apparent on the Provo eyetracking (FP and GP) measures, where cloze surprisal explained reading times over and above the manipulated GPT2 variants. The authors speculate the effect is smaller on UCL eyetracking because that corpus consists of short, isolated sentences made up of high-frequency words, so the H3 manipulation would barely change predictions.

  • Similarity-adjusted surprisal underperformed. Both SA cloze surprisal and SA GPT2 surprisal were poor predictors of reading times, indicating that simple "count-and-divide" is not an unreasonable way to convert cloze responses into probabilities. Results were inconclusive about which set of alternatives (cloze responses vs. GPT2 samples) is better.

Methodology in Plain English

The authors assembled four English datasets that contain both cloze responses and reading times: BK21 SPR (216 sentence triplets, about 90 cloze responses per sentence, with target-word prediction rates of roughly 91%, 20%, and 1% across high-, moderate-, and low-cloze conditions, plus SPR times from 216 subjects), Provo ET (55 paragraphs totaling 2,746 words, cloze responses from 478 subjects at about 40 responses per word, and eye-tracking from 84 separate subjects), and the UCL SPR and ET datasets (361 sentences, SPR times from 117 subjects, a 205-sentence subset with eye-tracking from 48 subjects, and around 80 cloze responses per word from a separate study).

Reading times were filtered in standard ways: first and last words of each sentence and line were removed, SPR times and go-past durations above 3000 ms and first-pass durations above 2000 ms were removed, and trials with incorrect comprehension responses in the UCL dataset were dropped. The resulting data range from 40,993 observations (UCL GP) to 105,958 (Provo FP).

Cloze probabilities were smoothed so that words never produced by any participant still received nonzero probability, then compared across different ways of relating probability to reading time. Predictability from the LM came from GPT2 (small), with whitespace probability folded into word probability to preserve a proper distribution over words. Regression models were linear mixed-effects models with baseline predictors of word length, sentence position, unigram surprisal (from KenLM estimated on about 6.5 billion words of OpenWebText), and whether the preceding word was fixated for the eyetracking measures. Model fit was assessed by 10-fold cross-validation on held-out log likelihood, with significance determined by paired permutation tests and Bonferroni correction.

For the second experiment, the authors intervened on GPT2 itself: downsampling it to cloze-level resolution, collapsing its vocabulary into semantic clusters, and restricting it to frequent tokens. If a hypothesis is right, degrading GPT2 in that particular way should erase its advantage over cloze. The third experiment replaced raw observed-word probability with a similarity-weighted average over the alternative completions each source produced, using normalized cosine distance between GPT2 token embeddings. BK21 was excluded from that analysis because it does not release raw cloze responses.

Why This Matters

The paper shifts the conversation from "LM surprisal fits better, so use it" to "here is why it fits better, and what that implies." The answer matters because the reason determines whether LM surprisal is capturing human-like prediction or something else. If the advantage comes mostly from resolution, then cloze studies with more responses could close much of the gap. If it comes from fine-grained semantic and frequency distinctions, a different question arises: whether human comprehenders make those same distinctions at all.

  • Psycholinguistic modeling: The finding that cloze probability is better modeled with a quadratic transform (S(w_t)^2) than a linear one is directly actionable for anyone fitting reading time models with cloze norms.

  • Stimulus norming: The resolution result implies that cloze norming studies with more responses per item could produce stronger predictors, which affects how researchers budget and design norming experiments.

  • Cognitive model evaluation: The work cautions against treating LM probabilities as a one-size-fits-all explanation of processing effort, and argues for models (such as predictive coding models) that specify links between specific stages of probabilistic inference and specific measures.

  • Benchmark interpretation: For NLP work that uses reading time or cloze data as evaluation targets for LMs, this paper clarifies that agreement between LM probabilities and human reading behavior does not by itself demonstrate human-like prediction.

Industry relevance: The paper does not report industrial applications or deployments. Its relevance to industry is indirect — for practitioners who use cloze-style or LM-based probability estimates as signals in evaluation pipelines, and for anyone reasoning about what LM probability differences actually encode about human expectations.

Future Directions

  • Higher-resolution cloze studies. Since matching GPT2's resolution to cloze's erased the advantage, the authors call for cloze studies that collect substantially more responses than are typically gathered — while noting that more responses alone will not fix the inherent limitation that cloze is an untimed production task.

  • Alternatives to the traditional cloze task. The authors propose timed versions of cloze, or a maze-like variant, to control for factors like conscious reflection that can influence untimed responses.

  • Testing human sensitivity to fine-grained distinctions. Experiments should test whether humans' expectations differentiate between semantically related words or between low-frequency words, potentially using stimuli informed by LM probabilities.

  • Better ways to combine cloze with LM estimates. The SA surprisal experiments produced poor fits and inconclusive comparisons, so the authors explicitly leave the exploration of methods for combining cloze responses with LM-based estimates to future work. They also note that their findings may not generalize beyond English-language models and native-English-speaker data, as they are not aware of cross-linguistic datasets of cloze responses aligned to reading time data.

Target Audience

Researchers in psycholinguistics and computational cognitive modeling who use predictability estimates to model reading behavior; NLP researchers interested in what language model probabilities reveal about human language processing; and methodologists designing cloze norming studies or evaluating LMs as cognitive models. The paper assumes familiarity with surprisal, mixed-effects modeling, and the cloze task, so it is less suited to readers without that background.

Authors’ abstract

How predictable a word is can be quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs).When used as predictors of processing effort, LM probabilities outperform probabilities derived from cloze data. However, it is important to establish that LM probabilities do so for the right reasons, since different predictors can lead to different scientific conclusions about the role of prediction in language comprehension. We present evidence for three hypotheses about the advantage of LM probabilities: not suffering from low resolution, distinguishing semantically similar words, and accurately assigning probabilities to low-frequency words. These results call for efforts to improve the resolution of cloze studies, coupled with experiments on whether human-like prediction is also as sensitive to the fine-grained distinctions made by LM probabilities.

Read the original paper