Skip to content
AI.info

Natural language processing

Summarization and Content Condensation

Design extractive and abstractive summaries for documents, conversations, and evidence sets while controlling omission, distortion, and unsupported detail.

By the end you can

Example

A good summary is selective rather than merely short

Different readers need different condensations of the same incident report. Nothing about the source decides which of them is right. The purpose does, and a system that has not been told the purpose is not compressing. It is guessing.

Everything that follows is measured. A national standards body fixed the length budgets and the no-addition rule in writing, decades ago. A national agency ran the benchmarks that scored condensation for twenty years. And two shipped consumer systems condensed real news into claims their sources never made, in public, with dates.

  • An executive wants cause, impact, current status, and decision needed.
  • An investigator needs chronology, disputed claims, sources, and unresolved evidence.
  • A customer needs a clear explanation, remedy, and next step.
  • A monitoring system needs changed facts since the previous update.
  • A search preview needs enough context to judge whether opening the document is worthwhile.
  • An accessibility summary may need simpler language without removing essential conditions.

Comparison

Summary types and their obligations

The requested output should specify audience, purpose, scope, and update behavior. Two of these five families are precise enough to have a public record attached. One was written into an agency benchmark. The other has a company post-mortem.

Update summarization is a defined task, not a label. NIST ran it as a scored evaluation: “Piloted in DUC 2007, the TAC 2008 Update Summarization task is to generate short (~100 words) fluent multi-document summaries of news articles under the assumption that the user has already read a set of earlier articles.” The guidelines describe a reader of whom NIST writes that “he would like to read summaries that only talk about what's new or different”. The test set was about 48 topics. Each topic held 20 AQUAINT-2 news articles, split into a chronologically earlier Set A and a later Set B of 10 documents each, and Set A is stipulated as already read. That single stipulation is what turns temporal alignment from a good intention into a scored requirement. A fact correctly extracted from Set B is still a failure if Set A already carried it. Content was not scored by n-gram overlap at all. It was scored against multiple human model summaries with the Pyramid method, which weights Summarization Content Units by how many model summaries express them. Nenkova and Passonneau built that method in 2004 on the premise “that no single best model summary for a collection of documents exists”.

The multi-document row's warning that source authority varies has a public failure attached to it. In May 2024 Google's AI Overviews told people to put glue on pizza, and the screenshots going round were not fakes. Google said so itself on 30 May: “But some odd, inaccurate or unhelpful AI Overviews certainly did show up.” — Elizabeth Reid, VP, Search at Google. The explanation is the uncomfortable one for anyone who treats faithfulness and correctness as the same property. Some overviews had condensed satirical and forum sources accurately. They were wrong because the source was: “Forums are often a great source of authentic, first-hand information, but in some cases can lead to less-than-helpful advice, like using glue to get cheese to stick to pizza.” The advice to eat a rock a day came out of the same mechanism — a data void, filled from whatever occupied it. Google reported more than a dozen technical fixes and disclosed a rate: “We found a content policy violation on less than one in every 7 million unique queries on which AI Overviews appeared.” BBC News, reporting the episode, recorded that Google described the results as “isolated examples”, and named both the “non-toxic glue” pizza answer and the one-rock-per-day answer. A multi-document summarizer that weights every source alike inherits the least reliable one it read.

FigureComparison · 5 columns

Extractive summary

Select sentences or spans from the source.

  • High source fidelity
  • Can be disjointed
  • Preserves source wording
  • Limited compression

Abstractive summary

Generate a new formulation of the source content.

  • Flexible compression
  • Can integrate information
  • Risk of unsupported claims
  • Needs factuality checks

Query-focused summary

Condense evidence relevant to a stated information need.

  • Task-specific selection
  • Depends on query interpretation
  • Can omit unrelated context
  • Useful for research and support

Update summary

Describe what changed relative to earlier information.

  • Suppresses repeated facts
  • Requires temporal alignment
  • Sensitive to version errors
  • Useful for ongoing events

Multi-document summary

Combine several sources with overlap or disagreement.

  • Needs deduplication
  • Attribution and conflict matter
  • Source authority varies
  • Chronology can be difficult

Visual

A content plan before wording

Separating selection from realization makes summary errors easier to diagnose. This is not the lesson's own advice, and it is not new. Condensation has had a written standard since before neural summarizers existed.

The standard is ANSI/NISO Z39.14-1997, Guidelines for Abstracts, reaffirmed in 2015. In thirty pages it performs most of steps 1 to 4. It separates informative from indicative abstracts — a scope decision taken before a word is written. It names the content elements to be selected: purpose, methodology, results and conclusions. It sets the length budget by document type rather than by taste: a maximum of 250 words for papers and articles, 100 words for notes and short communications, 30 words for editorials and letters, and 300 words for monographs and theses. And it states step 4's whole constraint in one sentence: “Do not include information or claims not contained in the document itself.”

The standard also governs the part of selection that a length budget cannot reach: “Retain the balance and emphasis of the original documents, except in a slanted abstract.” The exception is the interesting half. A condensation written for a declared purpose is permitted to re-weight its source — that is what query-focused summarization is. The standard's answer is that re-weighting must be declared rather than performed silently. The Oregon Department of Transportation distributes the same text under the earlier 2010 reaffirmation, with the same no-addition sentence and the same length table.

FigureProcess · 5 steps
  1. 1. Define audience and decision

    What should the reader know or do after the summary?

  2. 2. Identify source propositions

    Extract events, entities, relations, dates, attribution, and uncertainty.

  3. 3. Select and order content

    Apply scope, importance, chronology, query, and length constraints.

  4. 4. Realize concise language

    Write or select text without adding unsupported links or causes.

  5. 5. Verify against the source

    Check every claim, omission, attribution, number, and caveat.

Analogy

Packing for a short but demanding journey

A traveller packs a small bag for a specific trip. Space is limited, so the right contents depend on destination, weather, activity, and consequences of forgetting an item.

Items in a bag keep their meaning wherever they sit. Summary facts are related, and can change meaning when separated or reordered. Packing still shows why “shorter” has no value without a use case.

Compression is an allocation of limited attention under a declared purpose.

Key idea

A fluent compression can create a new claim

On 13 December 2024 the BBC complained to Apple. An Apple Intelligence notification summary had grouped three separate BBC News stories into a single alert. Two were condensed accurately — one on Syria, one on South Korea. The third reported that Luigi Mangione, the man arrested over the killing of UnitedHealthcare CEO Brian Thompson, had shot himself. The BBC had reported no such thing. Every input sentence was true. The false claim was manufactured in the joining, and it was delivered to lock screens under the BBC's own masthead. The same feature had already produced a “Netanyahu arrested” alert from a grouping of New York Times stories on 21 November.

Four days after the complaint, Reporters Without Borders asked for the feature's withdrawal outright. “AIs are probability machines, and facts can’t be decided by a roll of the dice. RSF calls on Apple to act responsibly by removing this feature.” That was Vincent Berthier, who heads the group's Technology and Journalism Desk, on 17 December 2024. Apple gave way on 16 January 2025. It disabled notification summaries for all news and entertainment apps in the iOS 18.3, iPadOS 18.3 and macOS Sequoia 15.3 developer previews, began italicising every remaining summary, and added a per-app opt-out from the Lock Screen.

The mechanism is unremarkable, which is the point. Merging “the valve failed” with “maintenance was delayed” adds a causal link that neither sentence carried. Pronoun changes, date normalization and combined statistics do the same thing quietly. Verification must therefore compare propositions, attribution, numbers, modality and temporal scope against the source. Surface overlap would have scored the Mangione alert as three faithful condensations. At the level of overlapping words, that is exactly what it was.

A summary is faithful only when its compressed relationships remain supported.

Case

Hallucination showed up in more than 70 per cent of single-sentence summaries

How often does a compression invent something? Somebody counted. Annotators marked every hallucinated span in the output of several neural abstractive summarization systems, and the first conclusion is the frequency: “intrinsic and extrinsic hallucinations happen frequently – in more than 70% of single-sentence summaries”. The second is that these additions are rarely lucky guesses — “over 90% of extrinsic hallucinations were erroneous”. Maynez and colleagues published that in 2020. Seven summaries in ten carried a claim the document had not made. Nine in ten of those additions were simply wrong.

An extrinsic hallucination is the interesting category. It could in principle be true and merely absent from the source — the summarizer supplying real world knowledge the document happened not to state. That is the charitable reading, and the annotation rejects it. Under 30 per cent of single-sentence summaries came through clean.

Figure

An extrinsic hallucination could in principle be true and simply absent from the source; nine times in ten it was not true either. Maynez, Narayan, Bohnet and McDonald (ACL 2020); Fabbri and co-authors (TACL 9).

Example

Summary failures worth labeling separately

Each failure family suggests a different data or system intervention, and two of them now have a counted rate rather than a definition.

BBC journalists spent a month reviewing answers from ChatGPT, Copilot, Gemini and Perplexity, each given access to BBC News articles. 51% of all the AI answers about the news carried significant issues of some form. Distortion was counted separately: 19% of the answers citing BBC content introduced factual errors in statements, numbers or dates. So was misattribution, and it is the figure to remember for any system that promises quotation: “13% of the quotes sourced from BBC articles were either altered or didn’t actually exist in that article.” The BBC published those numbers on 11 February 2025. Pete Archer's summary of the finding was that the assistants “can produce responses to questions about key news events that are distorted, factually incorrect or misleading”.

The rates move, and they do not move to zero. The EBU/BBC study re-measured the BBC subset eight months later: “The share of responses with significant issues of any type improved from 51% to 37%”. A fourteen-point improvement still leaves more than a third of answers as the case the review process has to catch.

  • Omission: a critical condition, risk, actor, or next step is missing.
  • Intrusion: the summary introduces a claim not supported by the source — the Mangione alert is one; more than 70% of single-sentence summaries carried one in the 2020 audit.
  • Distortion: a supported fact is altered in number, polarity, time, cause, or modality — 19% of answers citing BBC content introduced an error in a statement, number or date.
  • Misattribution: a claim is assigned to the wrong speaker, document, or authority — 13% of quotes sourced from BBC articles were altered or did not exist in the cited article.
  • Redundancy: scarce summary space repeats one idea while excluding another; the TAC 2008 update task penalises exactly this by stipulating that Set A has already been read.
  • Incoherence: selected facts are locally correct but their sequence or references are confusing.

ROUGE measures overlap, not summary truth

ROUGE compares n-gram or sequence overlap with reference summaries, and it remains useful for reproducible comparison. A valid paraphrase can score lower. A high-overlap summary can preserve a source error or omit a crucial fact. That is not a defect discovered later — it is the definition. Chin-Yew Lin introduced ROUGE-N, ROUGE-L, ROUGE-W and ROUGE-S in 2004 as counts of overlapping n-grams, word sequences and word pairs between a system summary and human reference summaries. Nothing in a count of shared units inspects attribution, causation or truth.

The field's standard benchmark took the metric at exactly that value. At NIST's Document Understanding Conference in 2004 the length budget was hard, and measured in bytes rather than words: “The maximum target length for very short summaries was 75 bytes. For short summaries it was 665 bytes.” Anything longer was truncated before scoring, with no credit for coming in under. And the scoring ran on one channel: “The summaries in tasks 1-4 were evaluated solely by means of ROUGE (ISI's Recall-Oriented Understudy for Gisting Evaluation, alias RED) automatic (n-gram) matching.” The official measures were 1-gram, 2-gram, 3-gram, 4-gram and longest common substring. A system could lead that table with 665 bytes of well-overlapping words in which every quotation was attributed to the wrong speaker. By TAC 2008 the same agency was scoring content manually with the Pyramid method, and computing ROUGE and BE scores for every submitted run alongside it, precisely to track how well the automatic measures correlated with the manual ones.

Complement overlap with semantic, source-grounded, question-based, factuality, human, and task-based evaluation; validate automated evaluators against domain judgments before scaling them.

Fabbri and colleagues took that second channel seriously in Transactions of the ACL. They “re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion”. They also “consistently benchmark 23 recent summarization models”. The work uses model outputs “along with expert and crowd-sourced human annotations”. The study exists because the question was open, in their own words. “The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress.”

Reference similarity is one measurement channel, not a certificate of faithful compression.

Steps

Evaluate a summarizer by content obligation

The evaluation should reveal what is lost or invented under each compression setting. Step 4 — score with multiple channels — is the step usually left as a list of nouns. Here is what it looks like carried out at full size.

In October 2025, 22 public service media organisations in 18 countries and 14 languages ran the same test at once, coordinated by the EBU and the BBC. Journalists — fourteen NPR editorial staff among them — rated responses from ChatGPT, Copilot, Gemini and Perplexity against a multi-criterion rubric: accuracy, sourcing, opinion versus fact, editorialization, and context. The methodology appendix records 2,709 of a potential 2,760 core responses actually evaluated. The headline sits in the foreword: “Overall, 45% of responses contained at least one significant issue of any type. Sourcing is the single biggest cause of significant issues (31%).” That foreword is signed by Pete Archer of the BBC and Jean Philip De Tender of the EBU.

Separate criteria are what make the result usable. 20% of responses had significant accuracy issues, and on that criterion the four assistants sat close together, between 18% and 22% — a difference no product decision should rest on. Sourcing is where they came apart. Gemini had a significant sourcing issue in 72% of its responses, while every other assistant stayed below 25%. One blended quality score would have averaged 72% against under 25% into a single unremarkable number, on the criterion that produced most of the failures in the whole study. The report's own conclusion is that “AI assistants are still not a reliable way to access and consume news”.

FigureProcess · 5 steps
  1. 1. Define required and optional content

    Use audience, query, length, risk, and source type.

  2. 2. Annotate propositions and attribution

    Record critical facts, uncertainty, numbers, and source ownership.

  3. 3. Generate at several budgets

    Test whether failures appear as compression increases.

  4. 4. Score with multiple channels

    Combine references, source support, human rubrics, and downstream tasks.

  5. 5. Review consequential failures

    Inspect omissions, intrusions, distortions, attribution, and chronology by slice.

Audit a meeting summary against the transcript

Take one meeting transcript and identify decisions, owners, deadlines, rejected proposals, uncertainty, and unresolved questions. Compare these with a generated summary at two length budgets. Borrow the budgets from a standard rather than inventing them: ANSI/NISO Z39.14 allows 250 words for a paper and 100 for a short communication, and the gap between those two is where selection starts to hurt.

Mark every summary proposition as supported, distorted, unsupported, or ambiguous. Then score the summary on the five criteria the EBU/BBC reviewers used — accuracy, sourcing, opinion versus fact, editorialization, context — rather than on one impression of quality. That separation is what let their study report 72% on one criterion and 20% on another instead of a single flat verdict. Check the quotations separately and verbatim, since 13% of the quotes drawn from BBC articles in the February 2025 study had been altered or did not exist. Then ask a participant whether the summary enables the intended follow-up without hiding disagreement.

A summary should be judged by the actions and understanding it enables, not only by resemblance to another summary.

Key takeaways