Skip to content
AI.info

Research

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

Overview Research area: Time-series forecasting augmented with large language models, retrieval-augmented generation (RAG), and self-reflective agent memory. Technical level: Intermediate. Readers nee

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting
arXiv
2610.04109
Published
2026-10-02
Authors
Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das, Chun-Liang Li, Nanyun Peng, Thomas Hartvigsen, Jinsung Yoon, Tomas Pfister

AI summary

Overview

  • Research area: Time-series forecasting augmented with large language models, retrieval-augmented generation (RAG), and self-reflective agent memory.
  • Technical level: Intermediate. Readers need some familiarity with forecasting benchmarks (MAE, MSE, Brier Score) and retrieval-augmented LLM pipelines, but the core idea is explained conceptually.
  • Scope: A single paper presenting SEER, a closed-loop framework that turns forecasting errors into textual updates of a retrieval memory and a causal knowledge base, evaluated on six event-driven forecasting domains against 15 baselines and four LLM backbones.

What This Paper Is About

Real-world time series are pushed around by outside events such as policy changes, supply disruptions, and earnings surprises, so forecasting from historical numbers alone is often not enough. Language models can retrieve relevant news, but the authors argue standard retrieval-augmented approaches suffer from high noise, missing signals, and no causal reasoning about how events actually move a series. SEER closes the loop instead: after each prediction it compares the forecast to the outcome and writes that lesson into two separate text memories that shape the next forecast.

Key Contributions

  1. A closed-loop event-conditioning paradigm. Instead of static retrieval-augmented generation, SEER formulates event conditioning as an error-driven process in which prediction residuals continuously refine how external knowledge is acquired, retrieved, and filtered.
  2. Two decoupled feedback mechanisms. A reflective retrieval memory (split into a retrieval component and a selection component) that refines future search queries and suppresses noise, and a persistent causal knowledge base that abstracts episodic failures into transferable domain dynamics.
  3. Strict chronological control. All knowledge base updates, reflective memories, and retrieved events are constrained to information available on or before the forecasting cut-off, using the fact-checking filtering mechanism of LEAF, which the authors say was verified through human expert audit. This is designed to prevent look-ahead bias and data leakage.
  4. Comprehensive evaluation. Experiments across six domains show SEER outperforming 15 time-series forecasting baselines and the backbone LLM with event RAG, plus model-agnostic gains across four LLM backbones and transfer to non-time-series, event-driven forecasting.

Main Findings

  • Broad outperformance on time-series tasks: Across the Memory, SSD, Stock, Weather, and Electricity domains, SEER outperforms all three baseline families (seven supervised deep learning models, six LLM-based forecasters, and zero-shot time-series foundation models TimesFM 3.0 and Chronos-2) as well as the backbone LLM (Claude 4.8 Opus) with event RAG.
  • Large MSE gains: SEER achieves a 50.2% MSE improvement over the best baseline on Memory 6m and a 16.4% MSE reduction on Electricity.
  • MSE gains exceed MAE gains: On Memory 6m the MSE reduction is 50.2% versus a 16.5% MAE reduction, which the authors interpret as evidence that SEER mitigates extreme errors during sudden, event-driven shifts.
  • Advantage holds at long horizons: At 9 months, where autoregressive signals typically decay, SEER reduces MSE by 24.3% on Memory and 19.6% on SSD compared with the best-performing traditional deep learning baselines.
  • Stock direction accuracy: SEER reaches 38.25% accuracy versus 36.92% for LLM (Event RAG) and 34.67% for the best time-series foundation model reported in that column (TimesFM 3.0).
  • Model-agnostic improvements in ablations: Across Claude Opus 4.8, Claude Sonnet 4.6, Gemini Flash 3.1, and Gemini Pro 3.1, SEER improves average performance by +10.5% for Claude Opus 4.8 and +11.4% for Gemini Flash 3.1 over basic event retrieval. Against the no-event LLM, Claude Opus 4.8 improves by +19.5% and Gemini Flash 3.1 by +18.8%.
  • Causal knowledge matters: The ablation shows memory-guided search and filtering alone (SEER w/o Knowledge) helps, but adding distilled causal knowledge systematically elevates accuracy further.
  • Reasoning quality judged higher: Using Grok-4.20-Reasoning as an LLM judge on four reasoning dimensions, SEER wins 55.6% of pairwise comparisons against LLM (Event RAG) with a net win rate of +38.7 across all backbones, and 48.8% with a net +31.3 against SEER w/o Knowledge.
  • Gains transfer beyond numeric forecasting: SEER also improves on Polymarket event resolution (Brier Score) and on Selective Stock Forecasting, where it chooses and predicts at most six of 50 equities per day.
  • Selective abstention drives stock gains: On 29 mutually selected stocks, SEER and LLM (Event RAG) perform similarly (51.7% versus 48.3%), but SEER instead targets 43 different tickers it judges more predictable, achieving 51.2% versus the baseline's 32.7% on those.
  • Interpretable mechanisms: In a memory price case study at the 2025-06-01 cut-off, SEER anticipates a price surge using distilled signals like "HBM cannibalization," a "step-wise multiplier decay," and "UPSIDE SQUEEZE," which numerical models and the event-RAG baseline miss.

Methodology in Plain English

At each time step, the system is given an entity description, a forecast horizon, and a look-back window of historical values. A frozen language model does the predicting, so nothing is learned by updating model weights; instead, the "learning" is the optimization of the text given to the model. That text has two parts: a pool of retrieved external events and a causal knowledge base of distilled domain rules.

For events, the agent runs two rounds of general targeted searches to build an initial event pool, then augments it with guided searches driven by what it has learned. A retrieval memory formulates the searches; a selection memory filters the results down to the ones that appear predictive. The causal knowledge base supplies reasoning support to the frozen model.

Once the real outcome arrives, the agent reflects on the whole trajectory — query, context, reasoning, prediction, and ground truth — and diagnoses why the prediction missed. Crucially, the diagnosis splits into two separate updates rather than one: one updates the retrieval and selection memory (was the evidence missing or noisy?), and the other updates the causal knowledge base (was the reasoning flawed?). A condensation operator merges similar entries and compresses redundancy when a memory approaches its text capacity.

The backbones tested are Gemini 3.1 Pro, Gemini 3.1 Flash, Claude 4.6 Sonnet, and Claude 4.8 Opus. The input context window is 14 days for Weather, Electricity, and Stock, and 9 months for component prices. The context space is shared within a domain across forecasting targets, but updated sequentially on one sample at a time, and the Stock and event forecasting domains are split into 6 and 3 subdomains respectively.

Why This Matters

Impact on research. The paper reframes event-augmented forecasting as an error-driven, closed-loop process rather than a one-shot retrieval problem, and separates retrieval failure from reasoning failure in the feedback signal. It also provides a rigorous chronology discipline (all updates indexed strictly up to time step t, with LEAF-based fact-checking filtering) for evaluating LLM-based forecasters on backtests without leakage. The result that a frozen model can be improved purely through evolving text context is relevant beyond forecasting, to any agent that must act in a non-stationary environment.

Real-world applications:

  • Component pricing and procurement: Forecasting monthly memory and SSD prices, where the paper's case study centers on a 64GB-DDR5 module and a regime shift driven by HBM production crowding out standard memory supply.
  • Energy and grid operations: Forecasting daily maximum and minimum electricity load across grid regions, as well as daily temperature extremes, for capacity and demand planning.
  • Financial decision support: Predicting next-day price direction for 50 large US equities through a selective forecasting setup that is allowed to abstain when signals are weak.
  • Prediction markets and event resolution: Estimating evolving probabilities of tech, politics, and economy outcomes (for example, Federal Reserve interest rate decisions) before they resolve.

Industry relevance. The approach requires no fine-tuning, works with multiple commercial LLM backbones, and produces natural-language rationales alongside numbers, which matters for auditing in regulated domains. Its explicit concern with look-ahead bias speaks directly to the practical problem of backtests that look good historically and fail live.

Future Directions

  • Reducing backbone dependence: Gemini Pro 3.1 showed the smallest gains (net +13.0 against LLM (Event RAG)) and was the only backbone where SEER w/o Knowledge did not beat event RAG (net -0.5), raising the question of which model properties make self-evolving context most effective.
  • Managing memory growth: The framework relies on a condensation operator to merge and compress entries when memory approaches its text capacity; how this compression affects long-run knowledge retention over very long horizons is an open question not resolved in the reported experiments.
  • Extending to other frequencies and domains: The evaluated tasks span daily, monthly, and discrete-resolution settings, but not higher-frequency intraday series, and the stock task explicitly assumes only a small predictable subset of movements exists.
  • Sensitivity to retrieval quality and cost: SEER depends on the LEAF fact-checking filtering mechanism for temporal authenticity, and the paper does not report computation or retrieval cost for the reflection loops, leaving the cost-benefit tradeoff of the closed loop unquantified.

Target Audience

Researchers and practitioners working on time-series forecasting, retrieval-augmented generation, LLM agents, and memory-augmented reasoning will get the most from this paper. It is also useful for quantitative analysts and engineers in finance, energy, and supply-chain forecasting who need models that react to exogenous events, and for evaluation-focused readers interested in leakage-free backtesting of language-model forecasters. Readers wanting full theoretical treatment of the memory update operators should note that detailed prompt templates are deferred to the paper's appendices.

Authors’ abstract

Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, and an inability to reason causally about event impacts. We propose SEER (Self-Evolving Event Reasoning and Retrieval), a closed-loop framework that dynamically optimizes event conditioning for time series forecasting. SEER translates prediction errors into two decoupled feedback mechanisms: (i) a reflective retrieval memory that refines subsequent search queries and filters spurious noise, and (ii) a persistent causal knowledge base that distills transferable domain dynamics. SEER enforces strict chronological boundaries across both event retrieval and reflection, preventing look-ahead bias and data leakage. Across six volatile time-series benchmarks, SEER consistently outperforms state-of-the-art time series foundation models and language model baselines.

Read the original paper