Skip to content
AI.info

Research

LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA

Overview Research area: Natural Language Processing — evaluation of long-document, abstractive question answering over narrative text (books). Technical level: Intermediate. The paper assumes familiar

LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
arXiv
2510.13494
Published
2025-10-15
Authors
Tommaso Bonomo, Luca Gioffré, Roberto Navigli

AI summary

Overview

  • Research area: Natural Language Processing — evaluation of long-document, abstractive question answering over narrative text (books).
  • Technical level: Intermediate. The paper assumes familiarity with QA benchmarks, n-gram metrics such as ROUGE-L and METEOR, and the LLM-as-a-judge paradigm, but its central arguments are readable without deep technical background.
  • Scope (one sentence): The paper builds LiteraryQA, a cleaned and validated literary subset of the NarrativeQA benchmark, then uses it to determine which automatic evaluation metric best agrees with human judgments and to benchmark seven long-context language models.

What This Paper Is About

NarrativeQA is the most widely used benchmark for question answering over long narrative documents, but it contains mismatched book–summary pairings, extraneous non-narrative text, duplicated questions, and malformed or invalid reference answers. These defects make it unclear whether low model performance reflects genuine difficulty, dataset noise, or inappropriate evaluation metrics. The paper's goal is to produce a cleaner, more homogeneous literary benchmark (LiteraryQA) and to establish, through a meta-evaluation, which automatic metric should be used to score systems on it.

Key Contributions

  1. A refined dataset. LiteraryQA is a human- and LLM-validated subset of NarrativeQA restricted to literary works, produced by a two-phase refinement pipeline (document-level and QA-level) and released publicly along with code.
  2. A documented refinement pipeline with reported validation. The authors describe the document filtering, text cleaning, question deduplication, question refinement and answer refinement steps, and report inter-annotator agreement and correction accuracy for the QA-level steps.
  3. A meta-evaluation of automatic metrics. The paper measures system-level Kendall's τ correlation between human judgments and n-gram-based metrics (EM, F1, ROUGE-L, METEOR), a neural metric (BERTScore), and three LLM-as-a-judge setups (GPT 4.1, Claude 3.7 Sonnet, Prometheus 2 7B), in both reference-based and summary-based settings.
  4. A benchmark of long-context LLMs. Seven instruction-finetuned models — five open-weight and two API-based — are evaluated on LiteraryQA in open-book, closed-book, and summary settings.

Main Findings

  • NarrativeQA is measurably noisy. Of the 177 books in the NarrativeQA test set, 8 were mismatched (4.5%), 20 were theatrical plays (11.3%), and 11 were non-narrative documents (6.2%), for 39 removed documents (22%), leaving 138 documents. The paper also reports 125 (1.2%) duplicate questions removed and 1608 samples (38%) modified by the end of the pipeline.
  • LiteraryQA test set sizes. The test set moves from 355 documents and 10,557 QA pairs in original NarrativeQA, to a filtered 138 documents and 4,223 QA pairs, to a final 138 documents and 3,785 QA pairs.
  • Documents were shortened. After text cleaning, documents averaged roughly 3K tokens shorter than the original NarrativeQA texts. Question and answer lengths remained close to the originals (questions 8.60 ± 3.30 tokens in NarrativeQA versus 8.62 ± 3.24 in LiteraryQA; answers 4.22 ± 3.63 versus 4.33 ± 4.07).
  • n-gram metrics correlate poorly with human judgment. On NarrativeQA, EM, F1, ROUGE-L and BERTScore all show near-zero or negative system-level Kendall's τ (EM 0.0325, F1 0.0328, ROUGE-L 0.0291, BERTScore -0.0477), while METEOR reaches 0.1519. On LiteraryQA, correlations improve slightly (METEOR 0.4444, EM 0.0614, ROUGE-L 0.0580, F1 0.0574, BERTScore 0.0677).
  • METEOR is the best n-gram-based option. The paper identifies METEOR as the metric to prefer among n-gram approaches, attributing this to its stemming and synonym-resolution features, and notes that it maintains stable scores across bins of prediction–reference length differences.
  • LLM-as-a-judge agrees best with humans, especially in the summary-based setting. On LiteraryQA's reference-based setting, Prometheus 2 7B reaches 0.4499, Claude 3.7 Sonnet 0.3651 and GPT 4.1 0.3282. In the summary-based setting, correlations rise sharply: Prometheus 2 7B reaches 0.6881, GPT 4.1 0.5593 and Sonnet 3.7 0.5243.
  • Prometheus 2 7B benefits most from the cleaner data. Its reference-based correlation rises by 25 percentage points from NarrativeQA (0.2195) to LiteraryQA (0.4499), overtaking the larger API-based judges.
  • Metric scores improve under open-book benchmarking. For Claude 3.5 Haiku, ROUGE-1 rises from 0.2208 on the original NarrativeQA to 0.2305 on the filtered version and 0.2655 on LiteraryQA, with similar monotonic improvements in ROUGE-2, ROUGE-L, METEOR and F1.
  • Rankings differ by metric. Three of four n-gram metrics (ROUGE-L, EM, F1) and BERTScore rank NExtLong-8B as best among the seven models, while METEOR identifies GLM-4-9B as best. Prometheus 2 judgments on a sample of predictions instead favor a closed-source model, Claude 3.5 Haiku.
  • Human agreement was strong. Inter-annotator agreement on the QA-level annotation reached an average Cohen's Kappa of 0.83 for the pipeline-quality assessment, and Kendall's τ of 0.7876 (LiteraryQA) and 0.8098 (NarrativeQA) for answer scoring.

Methodology in Plain English

The authors start from NarrativeQA and run a two-stage cleanup. In the document-level phase, they manually examine every test-set book to remove mismatched book–summary pairings, theatrical plays and non-narrative texts, then download the original HTML versions of the books and apply heuristic rules to strip HTML, Project Gutenberg headers and footers, and license text. Applying these rules to the training and validation splits is left to a Llama 3.1 8B Instruct classifier validated on the test set.

In the QA-level phase, they use Claude 3.5 Haiku for several steps: removing duplicate questions via a ROUGE-L similarity threshold, identifying and correcting malformed or ill-posed questions, and identifying and correcting malformed or invalid answers. Two of the authors then check the LLM's work on a subset of 20 test documents containing 583 QA pairs (15% of the set) — around 30 hours of annotation per annotator — and find that cases where the model edited both the question and the answer often produced compounding errors, so those samples (308) are excluded entirely.

To decide how to evaluate systems, the authors sample 500 QA pairs from LiteraryQA, run them through seven models, and collect human scores from two annotators. They then compute system-level Kendall's τ between each metric's ranking of systems and the human ranking, comparing n-gram metrics, BERTScore, and LLM judges in reference-based and summary-based variants. Finally, they evaluate the seven models on LiteraryQA under open-book, closed-book and summary conditions.

Why This Matters

  • Impact on research. The paper argues that reported progress on narrative QA may be partly obscured by dataset noise and unsuitable metrics, and shows that cleaning the data changes both absolute scores and system rankings. It also demonstrates that a small open-weight judge model (Prometheus 2 7B) can outperform larger API models when the reference answers are clean, which is a practical finding for anyone designing evaluation pipelines.
  • Real-world applications.
    • Benchmarking and selecting long-context models for tasks that require reasoning across an entire book rather than a short passage.
    • Educational tools that ask comprehension questions about full-length literary works and need reliable automatic grading.
    • Publishing, audiobook and media workflows where summaries, synopses or study guides must be checked against long source texts.
    • Internal model evaluation for teams that cannot afford to run human annotation at the scale the paper describes, since summary-based LLM judging correlates best with human rankings.
  • Industry relevance. The paper quantifies how unreliable legacy benchmarks can be, which matters to anyone reporting model capability numbers on NarrativeQA or benchmarks that include it (the paper lists ∞Bench, L-Eval, LongBench and HELMET). It also highlights a cost–accuracy tradeoff: LLM-as-a-judge correlates best with humans but the authors note it is costly at scale and lacks transparency.

Future Directions

  • Automated refinement of the training and validation splits. The authors apply the full manual pipeline only to the test set and defer LLM-based filtering of train and validation data to future work, calling the required manual effort prohibitive at scale.
  • Evaluation models specialized for narrative. The paper suggests fine-tuning judges specifically for narrative evaluation, possibly grounded in structured knowledge or an external knowledge base, to improve accuracy and reduce cost.
  • Retrieval-augmented generation. RAG was deliberately excluded because the authors wanted to test full-document comprehension; they note exploring RAG in this setting as an open direction while cautioning that retrieving small fragments can break narrative flow.
  • Broadening beyond literary works. Because LiteraryQA excludes movie scripts and theatrical plays to stay homogeneous, the authors state it should not be treated as representative of the full narrative landscape.

Target Audience

Researchers and engineers working on long-context language models, question answering benchmarks, and evaluation methodology will get the most from this paper, particularly those who report NarrativeQA numbers or who need to choose an automatic metric for abstractive, long-document QA. It is also useful for practitioners building LLM-as-a-judge pipelines, since it compares a small fine-tuned judge against much larger API models and reports where each succeeds. Readers looking for a new narrative QA dataset rather than a critique and repair of an existing one, or for comparisons with retrieval-augmented methods, will find those aspects out of scope.

Authors’ abstract

Question Answering (QA) on narrative text poses a unique challenge to current systems, requiring a deep understanding of long, complex documents. However, the reliability of NarrativeQA, the most widely used benchmark in this domain, is hindered by noisy documents and flawed QA pairs. In this work, we introduce LiteraryQA, a high-quality subset of NarrativeQA focused on literary works. Using a human- and LLM-validated pipeline, we identify and correct low-quality QA samples while removing extraneous text from source documents. We then carry out a meta-evaluation of automatic metrics to clarify how systems should be evaluated on LiteraryQA. This analysis reveals that all n-gram-based metrics have a low system-level correlation to human judgment, while LLM-as-a-Judge evaluations, even with small open-weight models, can strongly agree with the ranking identified by humans. Finally, we benchmark a set of long-context LLMs on LiteraryQA. We release our code and data at https://github.com/SapienzaNLP/LiteraryQA.

Read the original paper