Research
Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
Overview Research area: Natural Language Processing — positional encodings, hierarchical document structure, and causal-intervention interpretability in Transformers (category cs.CL; arXiv:2609.23551v

- arXiv
- 2609.23551
- Published
- 2026-09-20
- Authors
- Shuyang Xiang
AI summary
Overview
- Research area: Natural Language Processing — positional encodings, hierarchical document structure, and causal-intervention interpretability in Transformers (category cs.CL; arXiv:2609.23551v2, dated 27 Sep 2026; single author: Shuyang Xiang; license CC BY 4.0).
- Technical level: Advanced. The paper assumes familiarity with rotary positional encodings, attention-weight analysis, bootstrap resampling, and causal mediation/patching methodology.
- Scope (one sentence): The paper asks whether a paragraph coordinate that is separate from reading order causally changes attention beyond token distance, and whether the depth of the resulting attention compression — rather than its mere presence — is the signature of genuine paragraph structure.
What This Paper Is About
Standard positional encodings treat position as a single reading-order coordinate, but reading order cannot distinguish two tokens that are equally far apart in the same sentence, in different sentences of one paragraph, or in different paragraphs. The paper builds a hierarchical rotary positional encoding (hRoPE) with separate paragraph, sentence, and token channels, then holds the token sequence completely fixed and intervenes only on the paragraph coordinate to see what changes. The goal is to determine whether hierarchical paragraph position is a real causal factor in attention, and whether the depth of the resulting compression tracks true paragraph structure rather than any density-matched coordinate.
Key Contributions
- Precondition (Section 4): Reading order is shown to be an insufficient coordinate for hierarchical structure — an n-token sequence admits 2^(n−1) paragraph segmentations — and intervening on the paragraph coordinate p₁ with tokens held fixed causally changes attention in all three corpora.
- Characterization (Section 5): Attention is compressed near paragraph boundaries in every corpus, but compression alone is not diagnostic. An architecturally identical model whose p₁ channel carries density-matched, per-step-resampled random labels is also compressed, so depth, not location, separates real from random structure.
- Comparison with corpus-only candidates (Section 6): None of eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence) reproduces the cross-corpus ordering of compression depth, though embedding-based coherence comes closest; at the paragraph level the coherence relation transfers as a common slope with the opposite sign to the corpus-level ranking.
- A reusable measurement protocol: A token-distance-exact estimator (contrasting each paragraph displacement against the same-distance baseline, requiring ≥2,000 token pairs per cell, d ≤ 1,024) with paired document-cluster bootstrapping and anchor-normalized boundary-effect statistics.
Main Findings
- Merging compresses, splitting dilates. In all three corpora, fake-merge (giving two genuine paragraphs the same p₁) produces a negative effect and fake-split (inserting an artificial boundary inside a paragraph) a positive one. Table 1: WikiText-2 Δ_B = −1.300 (−1.395, −1.211) and Δ_C = +0.841 (0.723, 0.964); OpenWebText Δ_B = −1.302 (−1.372, −1.240) and Δ_C = +0.914 (0.835, 0.997); Code Δ_B = −3.769 (−4.526, −3.234) and Δ_C = +3.373 (2.800, 4.148). Brackets are paired document-cluster bootstrap 95% CIs, three seeds pooled per corpus.
- The manipulation is quantitatively consistent with natural boundaries. The normalized fake-split response β_C is 1.060 (0.985, 1.142) for WikiText-2, 1.069 (1.019, 1.122) for OpenWebText, and 1.522 (1.312, 1.797) for Code. The interval excludes the β_C = 1 calibration point for Code (significantly stronger than the average real boundary) and marginally for OpenWebText, while including 1 for WikiText-2. Percentile ranks C_pct place the fake-split response near the middle of the real boundary distribution: 57.9% (51.7, 65.7), 58.9% (54.9, 62.6), and 62.4% (58.3, 66.7) respectively.
- Compression overshoots the within-paragraph baseline. Because Δ_B is measured against the real-boundary level, the merged response 1 + Δ_B overshoots β = 0 in every corpus: −0.300 for WikiText-2, −0.302 for OpenWebText, and −2.769 for Code, whose sentence-boundary level already sits far from the within-paragraph level (β_A^sent ≈ −1.85, versus ≈ +0.2 for WikiText-2 and OpenWebText).
- The channel's content matters, not just its presence. Substituting the paragraph coordinate on a trained hrope_axial checkpoint gives validation losses ordered real < period < rand < const0 in all three seeds of every corpus: Code 2.383 / 2.480 / 2.641 / 2.720; WikiText-2 5.411 / 5.479 / 5.562 / 5.758; OpenWebText 5.622 / 5.683 / 5.720 / 5.922. Collapsing p₁ to a constant is worst by 0.30–0.35 nats versus real. With three seeds the minimum one-sided sign-test p-value is 0.125, so this is directional evidence rather than a formal significance test.
- Low language-modeling cost. hrope_axial is at most ≈0.5% above flat in validation loss (Appendix D) and never worse than an architecturally identical channel carrying no true positional information. The real-versus-rand gap is largest in Code and smallest in OpenWebText (0.26, 0.15, 0.10 nats for Code, WikiText-2, OpenWebText), and a logistic-regression probe on the residual stream reproduces the same Code > WikiText-2 > OpenWebText gradient.
- Depth separates real from random in two of three corpora. Table 3: Code hrope U* = −0.772 vs rand U* = −0.238, gap [−0.578, −0.498], significant; WikiText-2 hrope −0.481 vs rand −0.220, gap [−0.341, −0.171], significant; OpenWebText hrope −0.305 vs rand −0.326, gap [−0.059, +0.093], not resolvable (CI includes 0).
- Real depth varies more across corpora than the control's. rand_axial's three depths span −0.22 to −0.33, while hrope_axial's span −0.31 to −0.77 — a difference of degree, since all six model–corpus cells are compression.
- Location is not identifiable; depth is. In all three corpora the fitted degree-2 curve has an interior vertex, but WikiText-2's vertex location is structurally non-identifiable (median paragraph length ≈ 122 tokens leaves too few boundaries for cells to pass the n ≥ 2,000 gate beyond Δp₁ ≈ 3), and Code's degree-2 (≈5.5) and degree-3 (≈3.4) fits disagree. Depth is bootstrap-stable, with WikiText-2's depth shifting by ≈0.1 across sampling gates. The cross-corpus depth ordering (deep to shallow: Code −0.772, WikiText-2 −0.481, OpenWebText −0.305) is treated as descriptive, not inferential.
- No corpus-only quantity reproduces the depth ordering. The structural persistence length λ_struct (bootstrap median) is 4.85 (3.80, 6.70) for Code, 3.74 (2.94, 5.09) for WikiText-2, and 3.85 (2.49, 8.19) for OpenWebText — an ordering of WikiText-2 < OpenWebText < Code, nominally the reverse of the depth ordering on the two prose corpora. The paper notes λ_struct gets only Code's rank right, its prose-corpus comparison is fragile (medians 0.11 apart, overlapping intervals, and point estimates 5.48 / 4.90 / 3.37 that order the corpora differently), and its 500/500 bootstrap replicates completed for the main analyses.
- Embedding coherence comes closest. Of the eight candidates in Table 4, only embedding coherence's aggregate value (Code 0.096, WikiText-2 0.066, OpenWebText 0.051) matches the ordering while separating any pair (2 of 3 pairs: Code from both other corpora, but not WikiText-2 from OpenWebText). Coherence half-life (2.53 / 2.47 / 2.18 paragraphs) matches directionally but separates 0 of 3 pairs. λ_local (2.22 / 2.16 / 1.38) matches the ordering but also separates 0 of 3 pairs. λ_struct, δ_1/2 (3.36 / 2.59 / 2.67), L_int (3.84 / 3.02 / 3.30), δ_zero (5.17 / 6.93 / 3.72), and median paragraph length (55 / 122 / 45 tokens) all fail to match. A vocabulary-robustness check (Appendix F, Table 14) confirms λ_struct's ordering is unchanged at every tested cutoff.
- At the paragraph level the coherence relation reverses sign and is shared. Controlling for length, lexical diversity, and position, paragraphs with more similar neighbors compress less deeply, with no detectable slope difference between any pair of corpora, while length, lexical-diversity, and position effects remain corpus-specific. For pairwise slope differences, |z| > 2.87 is significant after Bonferroni correction over twelve comparisons. The full coefficient table (Table 5) is truncated in the available text — only the Length row is visible (Code −0.044, p<0.001; WikiText-2 −0.016, p<0.05), so the remaining per-corpus coefficients are not reported in the content provided.
Methodology in Plain English
The researchers replace a single linear position number with three separate numbers per token: one for which paragraph, one for which sentence inside that paragraph, and one for which token inside that sentence. Each number drives its own block of rotary channels, so the paragraph number can be changed without touching anything else.
They train several otherwise identical models on three structurally different corpora — WikiText-2, OpenWebText, and Python source code — using a fixed architecture of 8 layers, 8 heads, d_model = 512, context length 1024, 5000 training steps, and three seeds per corpus. One model has no paragraph channel (flat), one has only sentence and token channels (sent_axial), one has the full hierarchy (hrope_axial), and one has a paragraph channel filled with density-matched random labels resampled at every training step (rand_axial). A fifth model, period_axial, is a mirror control whose paragraph coordinate is the mechanical grid floor(t/L); it appears only in the loss-based substitution check (Section 4.3) and in neither the depth comparison nor the validation-loss comparison of Appendix D.
To see whether paragraph position matters, they edit only the paragraph number of an input: fake-merge gives two real paragraphs the same number, fake-split splits one paragraph into two. Because the tokens and their reading order are untouched, any change in attention cannot be caused by word content. They then average attention over all layers and heads, take logs, fit and subtract a straight line in token distance — the RoPE attention score decays roughly exponentially with distance, so logging makes that leading trend linear — and study the residual, called D_eff. A normalized boundary effect β places each manipulated condition on a scale where 0 is the within-paragraph baseline and 1 is the average real paragraph-boundary effect, with uncertainty from paired document-cluster bootstraps.
For the depth question they regress the residualized attention on paragraph displacement while contrasting each displacement against the same-token-distance baseline and pooling by pair count, keeping only (distance, displacement) cells with at least 2,000 token pairs and capping d at 1,024. Compression means the fitted displacement response is negative; depth U* is the value at the curve's interior vertex. Finally, they compute a set of quantities from the raw text alone — lexical persistence, paragraph length, and embedding-based coherence using all-MiniLM-L6-v2 — and check whether any of them predicts the ordering of depth across the three corpora, then repeat the coherence test one level down at the individual paragraph.
Why This Matters
Impact on research: The paper argues that compression location is a red herring and compression depth is the reproducible signature, and it supplies a falsification-friendly protocol: because a density-matched random channel also compresses attention, studies that claim a structural effect from the existence of compression alone are, by this account, under-constrained. It also reframes a modeling question (how to encode hierarchy) as a causal question (does an explicit hierarchical coordinate change attention, and with what shape), and it shows a hierarchical coordinate can be added at little language-modeling cost — at most ≈0.5% above flat in validation loss in this setup.
Real-world applications (note: the paper reports no deployed systems, products, or benchmarks; these are implications a reader might draw, not results the paper establishes):
- Long-document retrieval and summarization, where respecting genuine paragraph boundaries rather than fixed windows could change which passages attend to which.
- Code assistants, since Code showed both the deepest compression (U* = −0.772) and the largest real-versus-random loss gap (0.26 nats), suggesting block structure in source files is a distinct and strongly tracked signal.
- Evaluation and auditing of positional-encoding changes, using the density-matched random control as a falsification baseline before claiming a structural benefit.
- Chunking strategies in retrieval-augmented pipelines, where paragraph-level coherence (the candidate that came closest to matching depth) is the corpus statistic most worth monitoring.
Industry relevance: The results speak to a practical question for anyone pretraining or fine-tuning on long documents: does the model benefit from being told where paragraphs are, or does it merely follow any smooth density-matched coordinate it is given? The paper's answer in this setting is that the trained weights discriminate on content (real < period < rand < const0 in every seed), which is the condition under which an explicit paragraph channel is meaningful rather than decorative.
Future Directions
- Resolve the OpenWebText null. The hrope − rand depth gap there is [−0.059, +0.093], and the paper states it cannot separate a true null from a power limitation at three seeds — more seeds or more data would settle this.
- Test more corpora. All cross-corpus orderings, of U* and of every candidate explanation, rest on three corpora and are described as anecdotal in the statistical sense; additional corpora were not run.
- Explain the sign reversal. The paragraph-level coherence relation is a common slope across corpora (no detectable pairwise slope differences) but has the opposite sign to the corpus-level ranking — a discrepancy the paper reports without a mechanism.
- Fix identifiability of the compression vertex. Δp₁* is structurally non-identifiable for WikiText-2 and fit-sensitive for Code (degree-2 ≈5.5 vs degree-3 ≈3.4), so the location-side story is left open and only depth is treated as primary.
- Extend beyond paragraph scale. The protocol deliberately does not extend to the sentence level, where Δp₁ and token distance are harder to disentangle; that level remains unaddressed.
Target Audience
Researchers working on positional encodings, long-document and hierarchical Transformer architectures, and mechanistic interpretability or causal-intervention methodology will get the most from this paper, since it assumes fluency in RoPE variants, attention-weight analysis, and bootstrap inference. It is also useful for practitioners deciding whether to add explicit document-structure signals to a pretraining pipeline, though such readers should note that the evidence is limited to three corpora, three seeds, and one architecture (8 layers, 8 heads, d_model = 512, context length 1024, 5000 steps).
Authors’ abstract
Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate p1, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed relative to a token-distance-matched baseline in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too, more shallowly. What distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.