Research
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding Overview Research area: Natural Language Processing — long-context understanding, evidence selection, context com

- arXiv
- 2609.31382
- Published
- 2026-09-25
- Authors
- Zhaoyuan Xia, Qinghongbing Xie, Yung Xiang Hue, Jianguang Jiang, Gaofeng Lu, Zhenyu Jiao, Xing Yuan, Dai Dai, Tong Mo, Long Zeng
AI summary
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context UnderstandingOverview
Research area: Natural Language Processing — long-context understanding, evidence selection, context compression, and reinforcement learning for large language models (LLMs).
Technical level: Intermediate. Readers should be comfortable with LLM fine-tuning (SFT), reinforcement learning from verifiable rewards (GRPO), retrieval, ROUGE-L and F1-style metrics, and benchmark evaluation. The paper's core idea, however, is intuitive: select the relevant parts of a long document, summarize them in a question-conditioned way, then answer.
Scope (one sentence): The paper proposes a compress-then-reason paradigm (H2S) that makes a model explicitly highlight source-addressable evidence and write a question-conditioned summary before answering, trained on a new 6,647-example dataset and optimized with a process-level reinforcement learning reward.
What This Paper Is About
Long documents, multi-turn conversations, and code repositories contain far more irrelevant text than useful evidence, and the evidence that matters is often sparse and scattered across distant positions. Existing methods mostly help a model find relevant passages, but leave the harder job — combining those scattered passages into something coherent — implicit inside final-answer generation, which encourages unreliable and unnecessarily long reasoning. H2S instead forces the model to externalize both steps in a single autoregressive trajectory: first produce traceable evidence spans tied to document block identifiers, then produce a compact, question-conditioned summary, and only then answer.
Key Contributions
- A compress-then-reason paradigm (H2S) that externalizes evidence localization and evidence integration as two observable intermediate representations — source-grounded evidence (
E_q) and a question-conditioned summary (S_q) — generated before the final answer in one autoregressive trajectory, with the original document kept available throughout (organizing rather than pruning context). - H2S-Dataset, a long-context training set of 6,647 structured Evidence–Summary–Answer examples from 11 benchmark families, split into 4,228 examples for supervised fine-tuning and 2,419 for reinforcement learning, with an average context length of 43.9K tokens.
- H2S-RL, a GRPO-based reinforcement learning method whose deterministic trajectory-level reward jointly evaluates final-answer correctness, evidence grounding and span coverage, summary quality (ROUGE-L), and output structure through a structure gate — with no online LLM judge.
- H2S-Bench, a seven-task evaluation suite, and trained models (H2S-7B, H2S-14B) that achieve the strongest overall result among the evaluated open-source models, plus a process-level metric (Evidence–Summary Quality, ESQ) that measures intermediate quality rather than final answers alone.
Main Findings
-
A small question-conditioned summary boosts accuracy substantially. On DocQA-RL-1.6K with Qwen3.5-Flash as the answer model, a 0.46K-token Claude Sonnet 4.6 summary (~3% of a 15.19K-token document) used alone raised accuracy from 72.6 to 93.0, and combined with the document to 96.7; a 1.04K-token self-summary raised accuracy from 72.6 to 85.9.
-
H2S-14B leads evaluated open-source models. Under a shared 128K input and 4K output budget, H2S-14B averages 32.60 on the seven-task suite, beating the larger QwenLong-L1-32B by 5.01 points and Qwen3.8-27B by 10.17 points. H2S-7B averages 28.56, exceeding QwenLong-L1-32B by 0.97 points with fewer than one quarter of its parameters.
-
The gains span task types. H2S-14B ranks among the top two open-source models on five of the seven benchmarks, covering citation-grounded answering, precise retrieval, distributed-information aggregation, and long-document summarization. Closed-source models still score higher overall (Claude Opus 5 at 43.13, GPT-5.5 at 42.65, GLM-5.2 at 42.40, Kimi-K2.5 at 41.62, DeepSeek-V4-Pro at 37.14).
-
Reinforcement learning adds consistent gains over SFT. H2S-RL raises the 7B average by 4.88 points (23.68 → 28.56) and the 14B average by 3.44 points (29.16 → 32.60). For 7B it improves LongCite, AA-LCR, and HELMET-Summ by 7.11, 7.51, and 6.16 points; for 14B by 3.38, 4.46, and 8.93 points.
-
Both intermediate stages are necessary. Ablating evidence lowers the average by 4.01 points (7B) and 4.25 (14B); ablating the summary is more damaging, with drops of 9.90 and 8.47 points; removing both causes the largest drops of 13.21 and 10.32 points.
-
H2S is far less dependent on long generation. Increasing the output budget from 4K to 16K raises H2S-7B by only 1.92 points and H2S-14B by only 0.96 points (most of the gain already at 8K), whereas Qwen3.8-27B rises from 22.43 to 43.72. The abstract frames this as H2S-14B retaining 97.1% of its 16K-budget performance with only a 4K budget.
-
Process quality improves, not just answers. H2S-14B records the highest ESQ among all evaluated models at 31.02 (followed by GLM-5.2 at 30.50), and H2S-7B reaches 26.60; H2S-RL raises ESQ from 25.22 to 26.60 for 7B and 29.93 to 31.02 for 14B.
-
Attention becomes more evidence-focused. In a LongCite case study over 318 source blocks, H2S-7B more than doubles the share of block-level attention assigned to the 16 annotated evidence blocks compared with Qwen2.5-7B-Instruct-1M, while attention over most non-evidence blocks stays comparatively diffuse. The authors explicitly label this a mechanism example, not a performance estimate.
-
Higher completion rates under a tight budget. At 4K output, 87.11% of H2S-14B generations are complete versus 41.17% for Qwen3.8-27B; at 16K these become 93.98% and 86.60%.
-
Stated limitation. H2S is less naturally suited to tasks requiring broadly distributed context coverage, and process-level evaluation depends on construction-time evidence and summary references.
Methodology in Plain English
Data construction. Starting from existing (document, question, answer) instances, the authors build supervision for the two missing intermediate steps. Each document is split at semantic boundaries into blocks with stable identifiers, then hybrid lexical and semantic retrieval (BM25 plus embeddings with RRF ranking) builds a high-recall candidate pool. The question is decomposed into atomic information needs, and relevant claims are extracted and bound to a BLOCK_ID and a supporting span, forming the evidence supervision. Unsupported and duplicate claims are removed, and the retained claims are organized by information need into the question-conditioned summary; the information needs act as a coverage scaffold so a fluent summary cannot silently omit part of the question. An instance is kept only if it passes three checks — faithfulness (block IDs and spans align with source), coverage (every information need is supported), and answerability (the answer can be derived from the question and summary) — with failed instances sent back for reconstruction. Retrieval is used only for offline data construction; at inference the model reads the full supplied document with no external retriever.
Training. H2S-7B and H2S-14B are initialized from Qwen2.5-7B-Instruct-1M and Qwen2.5-14B-Instruct-1M. Both use full-parameter SFT (one epoch, learning rate 1×10⁻⁵, max input 128K, max target 4K), then full-parameter GRPO (one epoch per curriculum stage, learning rate 1×10⁻⁶, eight sampled candidates per question, at most 2K generated tokens). A progressive context-length curriculum runs 32K → 64K → 128K. All stages use bf16, a cosine schedule, 5% warmup, gradient checkpointing, and ZeRO-3.
Reward design. The GRPO reward multiplies a structure gate by a weighted combination of answer quality and the harmonic mean of evidence quality and summary quality. Answer quality follows each task's native output contract (numerical tolerance, choice accuracy, normalized exact match or token F1, set F1 or NDCG, ROUGE-L, or joint content–citation quality). Evidence quality combines grounding (fraction of evidence items whose block identifiers and span locators resolve in the source) and span-level F1 against reference intervals. Summary quality is ROUGE-L against the reference summary after removing local evidence identifiers. The harmonic mean means a strong component cannot mask a failed one, and the structure gate zeroes out malformed trajectories.
Evaluation. All models receive the same system prompt, question, and block-structured document, with a maximum input of 128K tokens and an output budget of 4K tokens; scoring uses only the final answer under each benchmark's native metric, and unparsable outputs receive zero. The suite comprises 2,575 examples spanning DocFinQA (numerical reasoning), Frames (multi-document reasoning), LongCite (citation-grounded answering), MRCR (precise retrieval), AA-LCR (distributed-information aggregation), LongBenchV2 (general long-context understanding), and HELMET-Summ (long-document summarization).
Why This Matters
Impact on research. The paper argues that evidence selection alone is insufficient and that the under-supervised step is evidence integration. By factorizing generation into highlight, summarize, and answer terms, it makes the middle of the reasoning pipeline observable, supervisable, and measurable — which is precisely what the ESQ metric targets, and what the authors identify as a gap in benchmarks that score only final outputs. It also shows that process rewards can substitute for an online LLM judge.
Real-world applications:
- Long-document question answering — contracts, manuals, and reports where a handful of clauses decide the answer (CUAD, Frames, and LongBench-Pro are among the largest training sources).
- Citation-grounded answering — producing answers with traceable, source-addressable references, which is what LongCite measures and what the
BLOCK_IDplus span design supports. - Repository-level code understanding — tracing a failing API call across an interface definition, an implementation in another file, and a configuration change, an example the paper raises explicitly.
- Multi-turn dialogue and agent histories — keeping long interaction histories compact so a model reasons over retained information rather than re-reading everything.
- Cost- and latency-constrained deployment — H2S-14B loses only 0.96 points going from a 16K to a 4K output budget, which matters when generated tokens dominate serving cost.
Industry relevance. The work is aimed at practical long-context serving: it targets models in the 7B–14B range that outperform larger open baselines (QwenLong-L1-32B, Qwen3.8-27B) while demanding far shorter generations. Two of the ten authors are affiliated with Baidu Inc., and the paper states that code and H2S-Dataset will be released, which lowers the barrier to adopting the recipe.
Future Directions
- Adaptive compression objectives. The conclusion names more adaptive compression objectives as future work, addressing the stated limitation that H2S is less suited to tasks needing broadly distributed context coverage.
- Stronger process supervision. The authors call for stronger process supervision beyond the current reward, which mixes answer, grounding, span, summary, and structure signals.
- Reducing dependence on construction-time references. Process-level evaluation currently depends on construction-time evidence and summary references, so a more self-contained way to judge intermediate quality remains open.
- Generalizing the paradigm. Since H2S is validated on seven tasks at 7B and 14B, whether the same highlight-then-summarize trajectory transfers to other scales, domains, and contexts beyond the 128K training stages (examples longer than 128K are retained but unused) is an open question. Citation coverage and information-need coverage are also logged for diagnosis but are not reward terms — leaving room to fold them in.
Target Audience
Researchers and engineers working on long-context LLMs, retrieval-augmented generation, and reasoning-model training — particularly those interested in process-level reward design and in metrics that evaluate intermediate evidence handling rather than final answers. It is also relevant to practitioners building document-QA, contract-analysis, citation-grounded, or code-repository assistants who face real input-length and output-budget constraints. Readers without a background in reinforcement learning or benchmark-based evaluation will need to consult the appendices for training and metric details.
Authors’ abstract
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, we construct H2S-Dataset, comprising 6,647 examples from 11 benchmark families with an average context length of 43.9K tokens, and introduce H2S-RL, which provides process-level rewards for evidence selection and summary construction in addition to final-answer correctness. We evaluate on H2S-Bench, a seven-task long-context suite. Under a shared 128K input and 4K output budget, H2S-14B achieves an average score of 32.60, outperforming Qwen3.8-27B by 10.17 points and obtaining the strongest overall result among the evaluated open-source models. H2S-14B also achieves the highest Evidence-Summary Quality score and retains 97.1% of its 16K-budget performance with only a 4K output budget. These results show that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.