Skip to content
AI.info

Research

Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation

Overview Research area: Natural Language Processing, specifically dialogue act (DA) annotation and dialogue segmentation for educational (tutoring and classroom) transcripts. Technical level: Intermed

Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
arXiv
2601.12061
Published
2026-01-17
Authors
Jinsook Lee, Kirk Vanacore, Zhuqian Zhou, Bakhtawar Ahtisham, Jeanine Grutter, Rene F. Kizilcec

AI summary

Overview

Research area: Natural Language Processing, specifically dialogue act (DA) annotation and dialogue segmentation for educational (tutoring and classroom) transcripts.

Technical level: Intermediate. The paper combines conceptual arguments about annotation design with quantitative benchmark-style comparisons, and the core metrics are built from label distributions rather than specialized neural architectures.

Scope in one sentence: The paper proposes conditioning dialogue segmentation on an annotation codebook ("codebook-injected" segmentation) and evaluates four segmentation methods — two LLM-based and two coherence-based — across two educational dialogue datasets using new evaluation metrics that do not require gold segment boundaries.

What This Paper Is About

Dialogue act annotation usually assumes that communicative or pedagogical intent is localized to a single utterance or turn. In real teaching dialogue, however, an instructional "move" such as giving an explanation often spans many turns, so annotators may agree on what the instructor is doing while disagreeing about exactly where that action starts and ends — boundary misalignment that gets misread as low reliability. The authors' goal is to separate boundary placement from label assignment by treating segmentation as an explicit first step conditioned on the annotation codebook, and to evaluate segmenters without patiently hand-labeled gold segment boundaries.

Key Contributions

  1. A codebook-injected (DA-aware) segmentation process that introduces an intermediate span level into the annotation workflow, explicitly decoupling boundary placement from label assignment so that multi-utterance constructs can be labeled more faithfully.

  2. An evaluation of LLM-based segmenters using both generic topic-shift prompting and DA label guidance, benchmarked against established NLP dialogue segmentation algorithms, including a retrieval-augmented variant that injects DA cues from semantically matched examples and DA labels.

  3. Gold-label-free evaluation criteria for segmentation in the absence of reference boundaries, capturing within-segment consistency, adjacent-segment distinctiveness, and segment-level distributional agreement between human and AI raters.

  4. Release of code and data at https://github.com/National-Tutoring-Observatory/codebook-injected-segmentation.

Main Findings

  • DA-awareness helps LLM segmenters' internal coherence, but not the coherence baseline's. Across both datasets, injecting the codebook most consistently improved within-segment coherence for LLM-based segmenters. On CLASS-annotated, GPT-5 DA-aware achieved the lowest normalized entropy (0.286) and highest purity (0.570), outperforming its text-only counterpart and all non-LLM baselines. On TalkMoves, the strongest coherence came from Gemini-3-pro DA-aware (entropy 0.566, purity 0.676), also improving on the text-only LLM variant. Dial-Start + DA-aware slightly degraded both entropy and purity relative to Dial-Start on both datasets.

  • Boundary distinctiveness depends on method family and dataset. On CLASS-annotated, Dial-Start achieved the strongest adjacent distinctiveness (JS_adj = 0.545), exceeding both LLM variants and Dial-Start + DA-aware. On TalkMoves, the strongest adjacent distinctiveness came from Gemini-3-pro DA-aware (JS_adj = 0.478), with Dial-Start close behind (0.475).

  • Local boundary contrast favors different methods again. On CLASS-annotated, GPT-5 DA-aware produced the highest boundary change rate (BCR = 0.288). On TalkMoves, the highest BCR came from Dial-Start + DA-aware (0.524).

  • Granularity differences are modest. All methods produced a comparable number of segments per dialogue, and differences were modest relative to high within-method variance. DA-aware prompting led to a moderate increase in segment count for LLM models, particularly on TalkMoves, without necessarily causing over-segmentation.

  • A three-way trade-off: no segmenter dominates. No method simultaneously optimized coherence, boundary separation, and human–AI distributional agreement. Improvements in coherence did not reliably translate into stronger boundary separation, and gains in either often coincided with worse human–AI alignment. DA-aware prompting tended to increase human–AI JS divergence for LLMs (for example, GPT-5 DA-aware on CLASS-annotated: 0.449 vs. 0.424 for GPT-5; Gemini-3-pro DA-aware on TalkMoves: 0.505 vs. 0.480 for Gemini-3-pro).

  • Human–LLM divergence is not purely model error. The authors argue that a codebook-injected LLM applies the taxonomy more literally than human annotators, who rely on pragmatic judgment and contextual smoothing and implicitly tolerate boundary ambiguity to preserve conversational flow. Divergence can therefore surface latent ambiguities or underspecified boundary conventions in human-annotated data.

Reported segmentation results (Table 3), mean [95% CI], lower is better for H̄ and JS̄_HA, higher is better for Pūr, JS̄_adj, and BCR:

Dataset / Method K (SD) H̄ Pūr JS̄_adj BCR JS̄_HA
CLASS-annotated, GPT-5 4.90 (1.71) 0.349 0.546 0.447 0.222 0.424
CLASS-annotated, GPT-5 DA-aware 6.30 (1.97) 0.286 0.570 0.477 0.288 0.449
CLASS-annotated, Gemini-3-pro 4.47 (2.01) 0.384 0.528 0.447 0.237 0.407
CLASS-annotated, Gemini-3-pro DA-aware 4.53 (1.70) 0.391 0.531 0.435 0.267 0.411
CLASS-annotated, Dial-Start 4.60 (0.56) 0.303 0.564 0.545 0.208 0.459
CLASS-annotated, Dial-Start + DA-aware 4.50 (0.63) 0.319 0.561 0.515 0.253 0.484
TalkMoves, GPT-5 10.86 (4.87) 0.616 0.659 0.447 0.235 0.470
TalkMoves, GPT-5 DA-aware 12.54 (6.97) 0.609 0.664 0.470 0.222 0.489
TalkMoves, Gemini-3-pro 16.62 (9.09) 0.598 0.662 0.471 0.235 0.480
TalkMoves, Gemini-3-pro DA-aware 19.53 (10.43) 0.566 0.676 0.478 0.252 0.505
TalkMoves, Dial-Start 14.60 (7.27) 0.619 0.640 0.475 0.416 0.513
TalkMoves, Dial-Start + DA-aware 14.75 (7.18) 0.639 0.633 0.469 0.524 0.503

Methodology in Plain English

The task. Given an ordered dialogue of T utterances (with speaker identifiers and DA labels removed from the input), a segmenter predicts a set of boundary indices, where a boundary index j means a boundary falls after utterance u_j. Those boundaries produce K = |B| + 1 contiguous segments, each intended to correspond to a span where the tutor's pedagogical intent is relatively stable.

Two datasets. The open-source TalkMoves dataset of authentic K–12 mathematics classroom discourse (63 sessions, 31,263 utterances, average 496 utterances per session, six pedagogically-grounded talk moves organized into three higher-level dimensions: Learning Community, Content Knowledge, and Rigorous Thinking) and a custom CLASS-annotated dataset of chat-based secondary school math tutoring collected and de-identified by an online platform called Upchieve, which connects volunteer tutors with students attending predominantly low-income (Title I) schools in the United States (30 randomly sampled sessions, 1,881 utterances, average 63 utterances per session, four adapted CLASS Instructional Support subdimensions: Feedback Loops, Scaffolding, Building on Student Responses, and Encouragement and Affirmation). For the CLASS-annotated dataset, annotations were produced by one expert with formal CLASS Instructional Support training. For TalkMoves, the study focuses only on teacher talk moves, using a subset of 63 sessions.

An LLM as a second rater. The original utterance-level DA labels came from expert human annotators. The authors additionally annotated the same transcripts with GPT-5 via the LiteLLM API, producing two parallel rater-specific label sets per dataset (Human and AI). This enables evaluation of segment-level rater agreement after aggregating utterance labels into segment-level distributions.

Four segmentation methods, in two families, each with a text-only and a DA-aware variant.

  • LLM segmentation, generic prompting: GPT-5 or Gemini-3-pro reads the full dialogue and returns a JSON list of boundary indices (0-indexed turn numbers marking the last utterance of each segment) using fixed prompts and decoding settings.
  • LLM segmentation, DA codebook prompting: the same models receive the codebook (DA definitions) and are told to place boundaries when the pedagogical function changes. The prompt explicitly says not to label the dialogue with the constructs but to use them only to guide segmentation.
  • Unsupervised coherence segmentation: Dial-Start, which trains an utterance encoder with a contrastive neighboring-utterance objective and computes a boundary score from changes in adjacent similarity.
  • Coherence segmentation with DA-conditioned retrieval: the authors keep a memory of expert-annotated human DA labels from the same dataset (1.9k labels for TalkMoves, 301 for CLASS-annotated), retrieve the top-K_ret semantically similar labeled utterances per utterance by cosine similarity, aggregate their DA embeddings into a single vector via similarity-weighted averaging, and fuse it into the utterance representation as normalized h_i + α·r_i, with α balancing the two contributions. The fused representation replaces the original in Dial-Start's coherence computations.

New evaluation metrics that need no gold segments. Each segment is converted into a distribution over DA labels, weighted by segment length (w_k = |S_k|/T). Three families of measures are reported:

  • Within-segment consistency: normalized entropy (divided by log₂C, so values lie in [0,1]; lower is better) and purity, the share of the dominant DA label in the segment (higher is better).
  • Adjacent-segment distinctiveness: Jensen–Shannon divergence between adjacent segment DA distributions, weighted by an adjacent-pair weight (|S_k| + |S_{k+1}|)/2T (higher is better), and boundary change rate, the fraction of predicted boundaries that coincide with a change in the utterance-level DA label (higher is better).
  • Human–AI distributional agreement: length-weighted Jensen–Shannon divergence between the human and AI DA distribution within the same segment (lower is better). This complements inter-rater reliability as a measure.

Compute resources. Dial-Start and its variant were run on a workstation with two NVIDIA Quadro RTX 6000 GPUs (24 GB VRAM each); a single GPU was used, with models implemented in PyTorch 2.9.1 (CUDA 12.8).

Why This Matters

Impact on research. The work reframes segmentation from a preprocessing afterthought into a first-class modeling decision that shapes how constructs are operationalized. It also supplies a gold-label-free evaluation toolkit for a setting where reference boundaries are expensive and themselves ambiguous, and it argues against reporting segmentation quality as a single score. The paper notes that connections between segmentation choices and annotation reliability remain underexplored, especially in educational dialogue where annotation targets reflect interactional strategies or pedagogical intent rather than topic alone, and boundaries may appear without strong lexical cues.

Real-world applications:

  • Scaling annotation of instructional dialogue at institutions that want to measure teaching quality but cannot afford exhaustive expert boundary coding.
  • Building and auditing LLM-assisted annotation pipelines for dialogue act taxonomies, where boundary disagreements can deflate apparent reliability (the authors cite the "unitizing" problem in prior agreement work).
  • Using a codebook-guided LLM as a diagnostic tool to surface latent ambiguities, theoretical drift, or underspecified boundary conventions in existing human-annotated datasets.
  • Selecting segmenters for specific downstream analyses — for example, identifying extended instructional phases versus detecting fine-grained shifts in dialogue flow.

Industry relevance. Any organization annotating multi-utterance spans — education technology platforms, tutoring providers such as the Upchieve-style synchronous chat tutoring setting described, and teams deploying LLM-assisted labeling at scale — faces the same unit-of-analysis problem. The paper's finding that no single method wins on all criteria means practitioners must choose segmentation to match their downstream objective rather than adopting a default best performer.

Future Directions

  • Extending the evaluation beyond the two datasets, two LLMs, and the restricted set of prompting and retrieval designs used here, to test generalization to multi-party or multimodal settings.
  • Developing metrics that capture pedagogically meaningful shifts not reflected by the distributional criteria used, particularly shifts requiring domain expertise.
  • Addressing the sensitivity of results to the underlying DA taxonomies and labeling quality, and to systematic differences between human and LLM annotations, which can influence the agreement-based measures.
  • Improving evaluation in low-density labeling regimes: the CLASS-annotated dataset contains very few labeled utterances, which may limit the resolution of the metrics and disproportionately influence human–AI agreement measures compared with more densely labeled corpora like TalkMoves.

Target Audience

Researchers and practitioners working on dialogue act and discourse annotation, dialogue segmentation, and LLM-assisted data labeling will benefit most, as will education researchers and learning-science teams who code instructional moves in classroom or tutoring transcripts. The paper is also useful for applied machine learning engineers building human-in-the-loop annotation pipelines, since it offers concrete metrics and a clear argument for matching segmenter choice to downstream analysis goals. Readers need no specialized neural architecture background, but familiarity with distributional metrics such as entropy and Jensen–Shannon divergence helps.

Note: the supplied paper content is truncated in Appendix B, so the full text of the individual dialogue act definitions (beyond the beginning of the CLASS "Feedback Loops" subcategory) is not available here. Appendix-based hyperparameter values (including the specific α and K_ret values) and the qualitative example appendices are likewise not reported in the available content.

Authors’ abstract

Dialogue Act (DA) annotation typically treats communicative or pedagogical intent as localized to individual utterances or turns. This leads annotators to agree on the underlying action while disagreeing on segment boundaries, reducing apparent reliability. We propose codebook-injected segmentation, which conditions boundary decisions on downstream annotation criteria, and evaluate LLM-based segmenters against standard and retrieval-augmented baselines. To assess these without gold labels, we introduce evaluation metrics for span consistency, distinctiveness, and human-AI distributional agreement. We found DA-awareness produces segments that are internally more consistent than text-only baselines. While LLMs excel at creating construct-consistent spans, coherence-based baselines remain superior at detecting global shifts in dialogue flow. Across two datasets, no single segmenter dominates. Improvements in within-segment coherence frequently trade off against boundary distinctiveness and human-AI distributional agreement. These results highlight segmentation as a consequential design choice that should be optimized for downstream objectives rather than a single performance score.

Read the original paper