Research
Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
Overview Research area: Multimodal dialogue and computational linguistics, specifically the study of common ground (grounding) in collaborative conversation using gaze behavior as observable evidence.
- arXiv
- 2609.18011
- Published
- 2026-09-16
- Authors
- Nan Li, Albert Gatt, Massimo Poesio
AI summary
Overview
Research area: Multimodal dialogue and computational linguistics, specifically the study of common ground (grounding) in collaborative conversation using gaze behavior as observable evidence.
Technical level: Intermediate. The paper's motivation and conclusions are accessible, but the methodology assumes familiarity with statistical hypothesis testing, effect sizes, clustered inference, and feature-based machine learning evaluation.
Scope: A cross-corpus empirical study that maps the gaze annotations of two collaborative-task corpora into a shared partner/task/away vocabulary and tests whether gaze patterns align with markers of mutual understanding.
What This Paper Is About
In tasks where two people hold different pieces of information, understanding cannot be assumed and must be actively built through interaction. This paper asks whether where people look provides usable evidence that this grounding process is succeeding. The authors compare two very different datasets: HCRC MapTask, where a giver directs a follower across mismatched maps, and MUNDEX, where an explainer teaches a board game to an explainee, to see whether gaze behaves consistently with respect to understanding across both.
Key Contributions
-
A shared gaze representation. A partner/task/away vocabulary that maps two distinct gaze annotation schemes (MapTask's up/down/off and MUNDEX's EX/EE/TABLE/AWAY) into one comparable space, designed to be reusable for other video-coded gaze corpora.
-
Cross-corpus evidence of directional convergence. Aligned reference interpretations in MapTask and "understood" (UND) judgments in MUNDEX are both associated with more task-directed gaze, less partner-directed gaze, lower gaze entropy, and fewer gaze transitions—despite the two corpora measuring related but distinct grounding constructs.
-
A within-speaker reference-chain analysis. In MapTask, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned, using 189 same-speaker pairs across 45 dialogues.
-
A role-stratified analysis and prediction benchmark. Associations concentrate in the participant leading the task (giver-produced references, explainer judgments), and a logistic-regression ablation shows modest recoverable signal over controls, with different gaze feature groups favored in each corpus.
Main Findings
-
Direction of effect is consistent across corpora. Aligned/UND windows show higher task-gaze proportions and lower partner-gaze proportions than non-aligned/non-UND windows. The largest pooled effect sizes are small: rank-biserial |r| = .058 in MapTask and .181 in MUNDEX.
-
Lower gaze entropy and fewer transitions signal understanding. Fewer label switches and less scattered gaze accompany positive grounding states in both corpora.
-
Effects concentrate in the task-leading role. In MapTask, associations are clearest for giver-produced references (six significant features, largest |r| = .086); follower-produced references show near-zero effects (largest |r| = .043). In MUNDEX, the explainer's judgments yield the largest effects (|r| = .206) and are the only stratum where features survive correction.
-
The explainee's gaze co-varies with the explainer's judgment. Explainee gaze proportions, entropy, and transitions correlate with UND in explainer judgments, indicating cross-participant coupling.
-
Gaze entropy drops at the moment of alignment. Within same-speaker MapTask reference chains, speaker entropy decreases significantly after correction (d_z = −.20, q = .044). This result is fragile: it does not survive correction when pair differences are averaged within dialogues (q = .20).
-
Eye contact matters in MapTask. All 13 structured features have larger effect sizes in the eye-contact stratum (largest |r| = .078 vs. .058 pooled); none is significant in the no-eye-contact stratum, where partner-directed gaze is largely absent. The formal condition interaction is not significant, so this is reported as a stratum difference rather than tested moderation.
-
UND is the task-directed extreme, but intermediate classes do not order consistently. The four MUNDEX understanding levels (UND, PART_UND, NON_UND, MISUND) do not follow a monotonic gaze gradient.
-
Prediction gains are modest and partition-dependent. Structured+temporal features give the highest macro-F1 in MapTask (.532 vs. .472 controls-only); raw proportions win in MUNDEX (.564 vs. .544 controls-only). MapTask gains over controls range from .015 to .070 across reshuffled partitions; MUNDEX gains reach at most .027.
-
Clustered inference weakens several results. When recurring participants rather than dialogues are the unit of inference, no MapTask feature survives correction; in MUNDEX, explainer task and partner gaze and explainee entropy and transitions remain significant, while explainee gaze proportions and mutual gaze do not.
Methodology in Plain English
The authors did not run new experiments. They re-analyzed two existing annotated corpora.
Building a shared vocabulary. MapTask labels each participant's gaze as up (toward the partner when eye contact is possible), down (at the map), or off. MUNDEX labels gaze as directed at the interlocutor, the table, or away. The authors collapse both into partner, task, and away, discarding gaze events with non-positive duration and resolving temporal overlaps within each participant's stream.
Defining windows. For MapTask, each window spans a reference expression plus 1.5 seconds of following context, capturing the addressee's immediate reaction. For MUNDEX, each window spans an understanding annotation plus or minus 2 seconds. Windows where either participant has less than 30% gaze coverage are dropped.
Extracting features. Six feature groups are computed per window: raw proportions, structured features (adding coverage, transition count, and duration-weighted Shannon entropy), temporal dynamics (run counts, durations, switch rate, latency, dominant/first/last label), transition bigrams (direction of gaze switches), coordination (joint gaze states sampled at roughly 10 Hz), and derived ratios (partner/task ratio, engagement, task dominance, asymmetries).
Testing associations. Binary contrasts use Mann–Whitney U tests with rank-biserial correlations, stratified by role and eye-contact condition, and checked with cluster-robust generalized estimating equations to respect within-dialogue or within-participant dependence. p-values are Benjamini–Hochberg adjusted.
Process analysis. Repeated mentions of the same landmark are grouped into reference chains, allowing a comparison of the speaker's gaze at the last non-aligned mention versus the resolving aligned mention.
Prediction. Logistic regression with standardized features, balanced class weights, and grouped cross-validation (10-fold by dialogue for MapTask, 5-fold by explainer for MUNDEX), with each feature group ablated.
Why This Matters
Impact on research. Grounding theory has long held that interlocutors seek evidence of understanding, but most computational work studies this through language alone. This paper demonstrates that a lightweight, ontology-level gaze representation can be compared across corpora, offering a template for cross-corpus multimodal analysis without requiring eye-tracking hardware. It also provides an honest account of how fragile such effects are when the unit of statistical inference is varied—a methodological caution that applies broadly.
Real-world applications:
- Conversational agents and social robots that modulate their behavior based on whether a user seems to be following an explanation.
- Telepresence and remote collaboration tools that adapt turn-taking or add clarification prompts when a partner's gaze signals confusion.
- Tutoring and instructional systems that estimate learner comprehension from nonverbal cues during explanations.
- Meeting analytics and accessibility tools that surface moments of communicative breakdown in recorded interactions.
Industry relevance. The findings are relevant to HCI teams building gaze-aware interfaces, AR/VR collaboration platforms, and automotive or teleconferencing systems where camera-based gaze estimation is already available. The paper's own caution applies: effects are small and partition-dependent, so gaze should be one input among several—alongside lexical content, dialogue acts, and task state—rather than a standalone signal.
Future Directions
-
Finer-grained gaze referents and dialogue actions. The three-category scheme cannot tell which landmark or object a participant is looking at, nor what function a glance serves. Distinguishing task-general from task-specific gaze behavior requires richer annotation.
-
Role-conditioned modeling. The paper reports a stratum difference rather than a tested moderation because feature-by-role interactions do not survive correction. Formally modeling role as a conditioning variable remains open.
-
Broadening the evidence base. Two asymmetric, face-to-face tasks in English and German support directional convergence; other task structures, languages, and multimodal predictors need testing. Larger corpora with non-recurring participants would address the inference fragility that limits several results here.
-
Joint modeling of dynamic grounding. Combining gaze with lexical content, dialogue acts, and task state in a single predictive or process model could test whether gaze adds signal beyond what language already provides.
Target Audience
Researchers and graduate students in computational linguistics, dialogue systems, and multimodal interaction who are interested in grounding, nonverbal communication, or cross-corpus methodology. It is also useful for HCI practitioners evaluating gaze as a signal in interactive systems, and for psycholinguists seeking computational operationalizations of common ground. Some background in statistics and machine learning evaluation is helpful for reading the results, though the introduction and discussion are broadly accessible.
Authors’ abstract
In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee's gaze. In same-speaker MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX. Because effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, we treat gaze as one contributing cue to grounding, to be interpreted alongside task and dialogue context.