Skip to content
AI.info

Research

Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study

Overview Research area: Natural Language Processing / search-augmented LLM forecasting evaluation, specifically temporal (information) leakage in web retrieval. Technical level: Intermediate. The pape

Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study
arXiv
2602.00758
Published
2026-01-31
Authors
Ali El Lahib, Ying-Jieh Xia, Zehan Li, Yuxuan Wang, Xinyu Pi

AI summary

Overview

Research area: Natural Language Processing / search-augmented LLM forecasting evaluation, specifically temporal (information) leakage in web retrieval.

Technical level: Intermediate. The paper is readable without deep ML background, but assumes familiarity with retrieval-augmented pipelines, forecasting evaluation metrics (Brier score), and the concept of information cutoffs.

One-sentence scope: A systematic audit of Google Search's before: filter and DuckDuckGo's date-range filter across roughly 390 resolved Metaculus forecasting questions, showing that date filters routinely admit post-cutoff information that inflates apparent forecasting accuracy.

What This Paper Is About

Retrospective forecasting evaluates forecasting systems on questions whose outcomes are already known, and it relies on an information cutoff so the model only sees evidence that existed before prediction time. Most pipelines enforce that cutoff with a search-engine date filter, assuming that filtering by date excludes documents published or updated after the cutoff. The authors test that assumption directly and show it fails: date-filtered retrieval routinely returns pages containing post-cutoff facts, up to and including the resolved answer, which makes the evaluation invalid.

Key Contributions

  1. Leakage audit. A systematic audit of Google Search's before: filter across 393 resolved forecasting questions and DuckDuckGo's date-range filter across 389, covering 38,879 and 34,454 fetched URLs respectively.
  2. Downstream impact measurement. A forecasting experiment with gpt-oss-120b comparing Brier scores with and without leaked documents, quantifying how much leakage inflates apparent accuracy (Brier 0.10 versus 0.24 with leak-free documents).
  3. Leakage mechanism taxonomy. Identification of four recurring pathways by which post-cutoff information enters date-filtered results: updated article content, related-content modules, unreliable metadata, and absence-based signals.
  4. Released artifacts. The questions used, generated queries, retrieved URLs, and audit time period are released at https://github.com/theolivecode/WebDataLeakageAudit.

Main Findings

  • Leakage is the norm, not an edge case. At least one retrieved URL contains topical post-cutoff information for 98.5% of questions on Google and 98.2% on DuckDuckGo. A weak signal or stronger appears for 94.1% and 96.1%; a major signal (score ≥ 3) for 71.0% and 81.2%; and a document that directly reveals the answer (score 4) for 41.0% and 54.8%.

  • URL-level rates. 33.2% of Google URLs (12,903 of 38,879) and 34.5% of DuckDuckGo URLs (11,898 of 34,454) contain post-cutoff information.

  • Leakage inflates forecasting accuracy. On 93 binary questions from 2025 with at least one score-4 document, providing only leak-free documents (score 0) yields a mean Brier score of 0.242, matching the no-retrieval baseline of 0.244. Providing documents scoring 3–4 (averaging 4.8 sources) drops the mean Brier score to 0.108, and restricting to score-4 documents only (2.6 sources) gives 0.129. For reference, predicting 50% on every question yields 0.25.

  • Corroboration matters. The slightly better score for the 3–4 condition than for score-4-only is attributed to additional score-3 documents giving context that helps the model avoid overreacting to a single snippet.

  • Leakage declines with more recent cutoffs. URL-level post-cutoff rates are highest for 2021 cutoffs (Google 46.3%, 1,831/3,955; DuckDuckGo 47.1%, 1,932/4,100) and 2022 (46.5%, 3,115/6,703; 48.0%, 3,170/6,605), drop for 2023 (34.5%, 2,008/5,821; 31.4%, 1,822/5,811), and drop further for 2025 (26.6%, 5,949/22,400; 27.7%, 4,974/17,938). The dataset contained no questions opened in 2024.

  • Four leakage mechanisms. Direct page updates (e.g. a missile tracking database first published in 2017 but continuously updated through 2023); related-content leakage (a 2016 article whose main text is pre-cutoff but whose related-articles sidebar revealed a December 2023 ICBM launch); absence-based signals (a comprehensive 1951–2025 conflict timeline with no mention of a US-Iran war); and unreliable metadata (a page reporting a 2020 last-updated date while containing text about Finland joining NATO in 2023 and other 2024 facts, scoring 4).

  • Judge reliability. Two annotators labeled disjoint subsets totaling 134 documents with at least 19 examples per score level. Human–LLM agreement was 76.1% exact accuracy when combining scores 0 and 1, with 0.85 Quadratic Weighted Kappa and an F1 of 0.82 for score-4 (direct leakage) cases.

Methodology in Plain English

The team gathered 393 resolved forecasting questions from Metaculus tournaments, spanning resolution dates from 2021 to 2025. For each question they used an LLM to generate 10 to 20 search queries from the question title and background, retrieved roughly 100 unique URLs per question through Google's before: operator and DuckDuckGo's date-range filter (start date 2000-01-01), and fetched each page's content. The cutoff was set to the question's opening date on the forecasting platform; DuckDuckGo returned usable results for 389 of the 393 questions.

Long pages exceeding 7,680 tokens were chunked into 256-token segments, and Maximal Marginal Relevance with Qwen-0.6B embeddings (λ = 0.7) selected up to 30 of the most relevant and diverse chunks; shorter documents were passed in full.

To score leakage at scale, they built an LLM-as-a-Judge system using gpt-oss-120b at temperature 0.5. Each request supplied the question title, background, resolution criteria, resolved answer, cutoff date, webpage content, and a 0–4 leakage rubric. The rubric spans: 0, no post-cutoff information or irrelevant post-cutoff information; 1, topical but uninformative; 2, weak directional signal; 3, major signal or decisive evidence for a partial component; and 4, directly reveals the answer. Absence-based signals are capped at a score of 3 to avoid over-interpreting omissions. The judge outputs a JSON object with a post-cutoff flag, a leakage score, and reasoning.

For the downstream experiment, the team evaluated gpt-oss-120b on binary questions opened in 2025 that had at least one score-4 document, since those questions post-date the model's knowledge cutoff. They compared Brier scores across retrieval conditions built from Google-audited documents only.

Why This Matters

Impact on research: Retrospective forecasting is used precisely because it allows fast, large-scale iteration using already-resolved questions. If the retrieval step leaks the answer, published accuracy numbers for search-augmented forecasters are not measuring forecasting ability at all. The authors argue date-restricted search on these engines is insufficient for credible retrospective evaluation and recommend stronger safeguards or evaluation on frozen, time-stamped web snapshots. The failure mode is not limited to forecasting: any pipeline treating search-engine date filters as sufficient — claim verification, dynamic fact-checking, timeline summarization — is potentially exposed.

Real-world applications:

  • Benchmark design for LLM forecasting and search-augmented reasoning systems, where contamination invalidates leaderboards.
  • Fact-checking and claim-verification systems that must reason only from evidence available at a given time.
  • Retrieval-augmented question answering over historical or time-sensitive events.
  • Auditing internal evaluation pipelines at organizations that backtest models on past events.

Industry relevance: Any team that backtests a retrieval-augmented system against known outcomes — forecasting platforms, search providers, AI labs running temporal evaluations — has a direct interest in the leakage rates reported here and in the four leakage mechanisms, at least one of which (unreliable self-reported timestamps) can bypass pipelines that double-check by filtering on extracted publication dates.

Future Directions

  • Test mitigations. The authors diagnose and quantify the problem but do not experimentally evaluate mitigation strategies such as Wayback Machine retrieval or frozen snapshot databases, and explicitly call these comparisons valuable future work.
  • Broaden beyond two engines and one platform. The audit covers Google Search and DuckDuckGo as representative major engines and Metaculus as the question source; leakage prevalence and mechanisms may differ for other retrieval systems.
  • Decouple judge and forecaster. Both the leakage detection and the forecasting experiments use gpt-oss-120b, raising the possibility of shared interpretation biases; future audits would benefit from judges drawn from different models.
  • Improve long-document handling. The MMR-based pipeline may omit dispersed leakage signals in excluded chunks, so more complete document coverage could shift measured leakage.

Target Audience

Researchers and engineers building or evaluating search-augmented LLM systems, especially those working on forecasting benchmarks, retrieval-augmented generation, or time-sensitive question answering. It is also relevant to benchmark maintainers and evaluation leads who need to know whether their date-filtered retrieval setup is actually enforcing the information cutoffs they assume, and to anyone comparing retrospective evaluation against prospective or snapshot-based alternatives such as ForecastBench, FutureX, and FutureSearch.

Authors’ abstract

Search-engine date filters are widely used to enforce pre-cutoff retrieval in retrospective evaluations of search-augmented forecasters. We show this approach is unreliable across two major search engines: auditing Google Search's before: filter and DuckDuckGo's date-range filter, we find that at least one retrieved page contains major post-cutoff leakage for 71% of questions on Google and 81% on DuckDuckGo, and the answer is directly revealed for 41% and 55%, respectively. Using gpt-oss-120b to forecast with these leaky documents, we demonstrate inflated prediction accuracy (Brier score 0.10 vs. 0.24 with leak-free documents). We characterize recurring leakage mechanisms, including updated articles, related-content modules, unreliable metadata, and absence-based signals, and argue that date-restricted search on these engines is insufficient for credible retrospective evaluation. We recommend stronger retrieval safeguards or evaluation on frozen, time-stamped web snapshots.

Read the original paper