Research
SagaScale: A Realistic, Scalable, and High-Quality Long-Context Benchmark Built from Full-Length Novels
Overview Research area: Natural Language Processing — evaluation of long-context understanding in Large Language Models (LLMs), specifically the construction of realistic long-context benchmarks. Tech

- arXiv
- 2601.09723
- Published
- 2025-12-27
- Authors
- Guancheng Du, Yong Hu, Wenqing Wang, Yaming Yang, Jiaheng Gao
AI summary
Overview
- Research area: Natural Language Processing — evaluation of long-context understanding in Large Language Models (LLMs), specifically the construction of realistic long-context benchmarks.
- Technical level: Intermediate. Readers will benefit from familiarity with context windows, retrieval-augmented generation (RAG), and agentic workflows, though the paper's argument is accessible without deep technical background.
- Scope: The paper introduces SagaScale, a bilingual benchmark of 1,124 question-answer pairs built from 103 full-length novels, and evaluates three long-context approaches across 12 frontier LLMs.
What This Paper Is About
Existing long-context benchmarks face a fundamental conflict among three properties: task realism, data scalability, and data quality. Synthetic benchmarks (e.g., RULER, Counting-Stars) are scalable but unrealistic; human-annotated benchmarks (e.g., NovelQA, LongBench-v2) are realistic but unscalable; and automated chunk-based generation (e.g., LaRA) is scalable but produces simpler, locally focused questions with potential factual errors. SagaScale addresses this by using an automated pipeline that supplies external resources such as Wikipedia pages to an LLM during benchmark construction — but not during evaluation — allowing the construction model to pose questions far more complex than any model can answer when given only the novel text.
Key Contributions
- Dataset and pipeline: A novel, automated data collection pipeline that leverages external resources to generate QA pairs, intended as a foundation for at-scale evaluation and training-set construction.
- Benchmark: SagaScale, a realistic, scalable, high-quality bilingual (English and Chinese) long-context benchmark featuring one of the largest context lengths to date, with average token counts exceeding 250K for English novels and 320K for Chinese novels. The comparison table reports a median context length of 228K.
- Analysis: An evaluation of three representative long-context approaches (Long Context, Naïve RAG, Agentic RAG) across 12 frontier LLMs, with in-depth analysis of their performance.
- Release: The SagaScale benchmark and the data collection codebase are publicly released at https://github.com/ducatiqwq/SagaScale.
Main Findings
- Direct full-context input can outperform retrieval workflows — when it fits: Restricted to problems that fit within each model's context window, Long Context surpasses Naïve RAG and Agentic RAG by 10.9% and 4.4% on average, respectively. For 1M-context models (GPT-4.1 and Gemini-2.5 series), Long Context performance consistently surpasses even the highest Agentic RAG score (64.8%).
- Overall averages across all problems: Naïve RAG 48.1%, Agentic RAG 55.8%, Long Context 27.7%. Long Context scores are depressed because processing failures for problems exceeding a model's maximum context length are recorded as failures rather than truncated.
- Agentic RAG beats Naïve RAG for nearly every model: Agentic RAG scores range from 40.3% to 64.8%, versus a tightly clustered 43.5% to 52.3% for Naïve RAG. The top Agentic RAG score is 64.8% (Claude-Opus-4-20250514, reasoning variant).
- Gemini-2.5-Pro is an exception among long-context models: It achieves 79.6% on Long Context and maintains stable accuracy across all book lengths, while GPT-4.1, GPT-4.1-mini, and Gemini-2.5-Flash show clear downward trends as length increases.
- Agentic RAG removes the Naïve RAG retrieval bottleneck: 66.7% of Naïve RAG instances yield either 0% or 100% accuracy across all models, versus only 25.4% for Agentic RAG. 72.6% of problems entirely unsolvable by Naïve RAG are solved by at least one model using Agentic RAG.
- All methods degrade with longer books, with Agentic RAG degrading least: Comparing cumulative accuracy up to 128K context versus up to 1M context, the drop is 6.8% for Naïve RAG, 5.5% for Agentic RAG, and 6.3% for Long Context.
- Formatting instruction adherence collapses in long contexts: Gemini-2.5 and GPT-4.1 maintain near-perfect adherence to the
boxed{}answer format, whereas models such as o1 and Claude-Sonnet-4 frequently fail — even though formatting requires only recalling a simple instruction. - Calibration error is detrimental for Agentic RAG but not for Naïve RAG: Naïve RAG calibration error rate is negatively correlated with Naïve RAG performance (Spearman's rho = -0.5) but positively correlated with Agentic RAG performance (rho = 0.6). The paper argues overconfidence is penalized in Agentic RAG because it causes premature termination of retrieval, while in Naïve RAG it can lead to occasional lucky guesses.
- Agentic RAG accuracy falls as tool-call chains lengthen: Accuracy declines across all models, including Claude-Opus-4 and DeepSeek-V3, as the number of tool calls k increases from 1 to 8.
- Closed-book accuracy is low, indicating limited contamination: The highest accuracy achieved by LLMs without access to the books is only 11.7%.
- Manual quality check: Of 100 randomly sampled instances verified by the authors, all were answerable and objective, and 95 out of 100 were fully correct.
- Question type diversity: Questions are well distributed across ten narrative-theory categories: Miscellaneous, Method, Quotation, Event, Entity, Character, Reason, Quantity, Location, and Temporal.
Methodology in Plain English
The pipeline has three stages.
Stage 1 — Book Collection. The authors collected 103 full-length novels: 77 in English and 26 in Chinese. Public-domain books came primarily from Project Gutenberg, and copyrighted books were obtained from the web. Each book had to satisfy four criteria: it must be a fictional narrative (to reduce contamination risk), typically exceed 100K tokens, be a single unified work rather than an anthology, and have an informative Wikipedia or Baidu Baike entry.
Stage 2 — External Resource Collection. Each novel is paired with its Wikipedia (or Baidu Baike) page, which is parsed into self-contained segments such as the "Synopsis" section. An LLM discards any segment containing extrinsic information (for example, adaptations) or subjective interpretations (for example, criticism), so that questions derived from those segments remain answerable from the book text alone.
Stage 3 — QA Curation. This splits into generation and filtering. For generation, the authors use a multi-query retrieval approach: an LLM breaks each page segment into concise sub-queries, each sub-query retrieves the three most relevant book chunks, and each chunk list is concatenated with the original page segment to form a two-granularity context (global information from the segment, local evidence from the chunks). DeepSeek-R1 generates QA pairs from this context. Prompts enforce that questions concern only fictional elements, are objective with a unique answer, are multi-hop across passages, and have a short answer.
Filtering has three steps. Correctness Verification uses GPT-4o Search Preview with web search to keep only QA pairs with confirmed correct answers. Realism Assurance extracts keywords from each question, runs a Google search, and uses the total number of search results as a popularity proxy; on a per-book basis, only questions scoring at least 0.01% of the highest-scoring question for that book are kept. Contamination Filtering prompts GPT-4.1, DeepSeek-R1, Claude-Opus-4, and Gemini-2.5-Pro to answer every question without the books; if any model answers correctly, the question is removed.
Evaluation setup. Three methods are compared. Long Context concatenates the instruction, full book text, and question; if the input does not fit the context window, a failure is recorded with no truncation attempt. Naïve RAG splits each book into 512-token chunks embedded with mE5-large and retrieves the top 3 chunks by embedding similarity. Agentic RAG iteratively issues search queries using the same retrieval setup, capped at a maximum of 8 retrieval requests. Twelve LLMs are evaluated, including Qwen3-235B-A22B, DeepSeek-V3, DeepSeek-R1, GPT-4o, GPT-4.1, o1, o4-mini, Claude-4, and Gemini-2.5, with temperature set to 0.0. GPT-4o serves as the LLM-as-a-judge, validated against 100 manually annotated samples with a Cohen's Kappa of 0.92. All token counts use the tokenizer from DeepSeek-R1-0528.
Why This Matters
The paper argues that long-context benchmarks have been stuck between unrealistic synthetic tasks and unscalable human annotation, and that automated alternatives such as LaRA produce simpler questions and rely on manually tuned seed examples (over 75% of LaRA's 128K book reasoning questions start with "why"). SagaScale offers a construction method that scales without that quality compromise, and its evaluation yields practical guidance for system builders: supplying the full context directly can beat multi-step retrieval when the model can actually handle the length, and agentic retrieval resolves the bottleneck of static single-round retrieval.
Real-world applications implied by the work:
- Long-document question answering, where users need answers grounded in an entire book, report, or record rather than a retrieved excerpt.
- Search agents and iterative retrieval systems, where the paper's calibration finding — that overconfident models terminate retrieval too early — informs agent design.
- Code repository understanding and long video understanding (for example, movies), which the authors explicitly name as domains their core methodology can be adapted to.
- Model selection and evaluation, giving practitioners a benchmark that separates genuine long-context comprehension from memorization.
Industry relevance: the paper's distinction between calibration behavior suited to single-pass tasks versus calibration behavior needed for strong search agents is directly relevant to teams building retrieval-augmented and agentic systems, and its release of both the benchmark and the full data pipeline supports reproducible comparison across frontier models.
Future Directions
- Scale the dataset for training purposes: SagaScale contains 1,124 QA pairs, which the authors describe as sufficient for benchmarking but insufficient for training, citing bottlenecks in manual novel/link collection and in the filtering process; they suggest that a single round of contamination filtering with one model may suffice for training data.
- Broaden the task scope beyond novel question answering, for example to large code repositories or long videos such as movies.
- Extend language coverage beyond the current English and Chinese bilingual setup, given the availability of novels and encyclopedic knowledge in other languages.
- Build more robust agents and workflows that maintain reliability over extended tool-call chains, since Agentic RAG accuracy declines as the number of tool calls grows.
- Investigate why Gemini-2.5-Pro maintains stable accuracy at extreme context lengths, which the authors note warrants further investigation.
Target Audience
This paper is most useful for researchers and engineers working on long-context LLM evaluation, retrieval-augmented generation, and agentic search systems. It also serves benchmark designers interested in automated, scalable data construction, and practitioners who need to choose between long-context models and retrieval workflows for document-heavy applications.
Authors’ abstract
Large Language Models (LLMs) have shown significant progress, but understanding long and complex documents remains challenging. Many long-context benchmarks have been proposed, but they face several limitations, including task realism, data scalability, and data quality. To this end, we introduce SagaScale, a realistic, scalable, and high-quality long-context benchmark built from full-length novels. The entire benchmark is constructed using an automated data collection pipeline that utilizes external resources (e.g., Wikipedia pages) to curate question-answer pairs. Critically, these external resources are provided only for benchmark construction and not during evaluation, which allows LLMs to curate complex questions that go beyond what they can answer during evaluation. SagaScale is also bilingual and offers the largest context length to date, with average token counts exceeding 250K for English novels and 320K for Chinese novels. Our evaluation across 12 frontier LLMs and three long-context methods -- Naïve RAG, Agentic RAG, and Long Context -- yields key insights, including: (1) Directly supplying the full context to the LLM can outperform other methods by a large margin; (2) Most LLMs still struggle with lengthy contexts, but Gemini-2.5-Pro stands out as an exception; and (3) Agentic RAG effectively addresses the retrieval bottleneck in Naïve RAG. Finally, we publicly release the SagaScale benchmark and our data collection codebase to facilitate future research.