Research
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
Overview Research area: Artificial intelligence evaluation, large language model capabilities, and open-domain forecasting (arXiv:2510.17638v2 [cs.AI], submitted 2025-10-20, revised 21 Dec 2025). Tech

- arXiv
- 2510.17638
- Published
- 2025-10-20
- Authors
- Qingchuan Yang, Simon Mahns, Sida Li, Anri Gu, Jibang Wu, Haifeng Xu
AI summary
Overview
Research area: Artificial intelligence evaluation, large language model capabilities, and open-domain forecasting (arXiv:2510.17638v2 [cs.AI], submitted 2025-10-20, revised 21 Dec 2025).
Technical level: Intermediate. The paper assumes familiarity with probabilistic forecasting concepts (Brier score, calibration, proper scoring rules) but explains each metric from first principles with worked examples.
Scope: The paper introduces and studies the "LLM-as-a-Prophet" paradigm — using LLMs to predict real-world future events — through a live benchmark called Prophet Arena that evaluates 23 models across 1,367 resolved events.
What This Paper Is About
The core problem is that despite rapid LLM progress, it remains unclear whether AI systems can reliably forecast open-domain real-world events — producing accurate predictions across many topics without domain-specific tuning. The authors build Prophet Arena, a continuously running benchmark that collects live prediction-market events, gives every model an identical information context, and scores forecasts along three separate dimensions rather than accuracy alone. Their goal is not only to rank models but to diagnose which underlying capabilities (knowledge recall, source interpretation, information aggregation speed, calibration) drive or limit predictive intelligence.
Key Contributions
-
Framing and naming a new paradigm. The paper defines "LLM-as-a-Prophet" — reliable open-domain prediction of real-world future events — as a distinct challenge for language model intelligence, and positions it as a contamination-free, objectively measurable testbed as established benchmarks approach saturation.
-
Building Prophet Arena. A live, extensible benchmark that adds 20 new events per day, decomposes forecasting into modular pipeline stages (event and market extraction, prediction context construction, probabilistic forecasting and evaluation), and supports a multi-horizon scheduling protocol. The authors state it is the first large-scale benchmark to run continuously on real-market events with a multi-horizon protocol inside a modularized pipeline.
-
Multi-dimensional evaluation. The paper argues that accuracy alone cannot measure probabilistic forecast quality (illustrated with a two-forecaster example) and evaluates models on three complementary axes: forecasting loss (Brier score), calibration error (ECE), and market return (Average Return), with market-implied probabilities as a built-in baseline forecaster.
-
A controlled empirical study. The authors evaluate 23 LLMs across 1,367 resolved events encompassing 72,136 markets, report five representative models in detail, and run mechanistic experiments on a uniformly sampled 100-event subset to probe how knowledge internalization, context construction, and reasoning shape forecasts.
Main Findings
-
LLMs are already competitive on Brier score, with a narrow spread. On the five highlighted models, Brier scores range from 0.184 (±0.006) for GPT-5^R to 0.219 (±0.008) for Llama-4-Scout, against 0.187 (±0.006) for the Market Baseline. The paper notes that all highlighted models fall in a [0.18, 0.22] band, where pure random guessing has an expected Brier score of 0.25.
-
Calibration differences are larger than loss differences. Strong models typically achieve ECE ≤ 0.05 while weaker ones fall in the [0.06, 0.2] range. Claude Sonnet 4^R leads on ECE (0.041), followed by GPT-5^R (0.042) and Grok-4^R (0.043); the Market Baseline sits at 0.069. All five selected LLMs calibrate better than the market baseline.
-
Rankings change depending on the metric. Model ordering by Brier score, ECE, and Average Return differ, which the authors present as evidence that the three dimensions capture genuinely complementary properties of a forecast.
-
LLMs do not reliably beat the market economically. GPT-5^R has the highest Average Return at 0.943 (±0.042) but still fails to reach break-even (Average Return < 1), and most highlighted models fall below 0.9. Returns carry wide confidence intervals because payoffs depend heavily on market-implied probabilities.
-
Extreme-bin calibration drives the gap between best and worst models. Reliability diagrams for OpenAI o3 (best) and the worst model by ECE show similar calibration in intermediate probability ranges, but o3 performs much better in the 0-0.1 and 0.9-1.0 bins, where it almost always predicts correctly — an advantage that appears frequently and helps explain gaps in both Brier score and market return.
-
Lead time changes the story, and markets learn faster near resolution. Average Brier score improves for both the market baseline and LLMs as resolution approaches. The market baseline lags several frontier LLMs at long horizons, but as resolution nears, markets incorporate breaking information faster and surpass LLMs in short-term accuracy.
-
Forecasts made too close to resolution must be excluded. Predictions within three hours of resolution are dominated by real-time information access rather than reasoning ability, and because Prophet Arena fixes the retrieval component across models, the authors exclude all such forecasts from evaluation.
-
The dataset is heavily sports-dominated. The 1,367 events resolved before October 11, 2025, collected via a real-time Kalshi data pipeline, break down as 76% Sports, 8% Entertainment, 7% Politics, and 9% Other. The authors verify that performance patterns are robust to reweighting toward a more balanced category distribution, reporting those results in Appendix B.8.
-
Identified bottlenecks. The abstract names three limits to superior predictive intelligence: inaccurate event recalls, misunderstanding of data sources, and slower information aggregation compared to markets when resolution nears. (The detailed experimental results supporting these, in Section 4, are not available in the truncated content provided.)
Methodology in Plain English
The researchers treat real prediction markets as a source of continuously generated, objectively scored forecasting questions. Each day they add 20 new events from active markets, and each event resolves to a verifiable Yes or No for every market underneath it. Because the events concern the future, no model can have seen the answer during training.
Every model receives the same package of information: web-sourced news titles, snippets, timestamps and URLs gathered by a single LLM-based search agent (GPT-4o with web access), plus market snapshots of latest Yes/No contract prices and trading volumes. Holding retrieval constant means differences in scores reflect reasoning and calibration rather than search skill.
Models output a probability between 0 and 1 for each market that it resolves Yes, plus a short written rationale. Only the probabilities go into quantitative scoring, and scoring happens only after outcomes are known. Three metrics are computed: Brier score (mean squared error between predicted probability and outcome, averaged within then across events), expected calibration error (comparing predicted probability to realized frequency, approximated with probability bins), and Average Return (a risk-neutral betting rule that allocates a $1 budget to whichever contract offers the better implied payoff, then measures the average payout across markets).
Market-implied probabilities serve as a synthetic baseline forecaster — for example, Yes and No prices of 0.8 and 0.2 yield an 80% baseline probability for Yes. This gives a difficulty anchor and a conceptually justified reference point for the human consensus. For the deeper mechanistic study, the authors use a 100-event uniformly sampled subset, released publicly.
Why This Matters
Impact on research. The paper argues that open-domain forecasting is a forward-looking alternative to benchmarks that are approaching saturation and are vulnerable to training-data contamination. By scoring forecasts on calibration and economic value alongside accuracy, it shows that a single metric can mislead: the same models reorder depending on which dimension is measured. It also supplies a baseline (market consensus) that earlier forecasting benchmarks typically lack.
Real-world applications:
- Finance and trading: Average Return directly measures whether a model's forecasts would be profitable against market prices, relevant to systematic trading and risk control.
- Policy and public decision-making: The paper cites forecasting's role in guiding high-stakes policy as motivation for reliable foresight.
- Risk-sensitive decision support: Because calibration captures the reliability of stated confidence, better-calibrated forecasts support better risk-preference-adjusted decisions even when absolute loss is slightly worse.
- Domain event monitoring: The benchmark spans politics, economics, sports, entertainment, and science, making it a template for tracking predictive performance in any area with verifiable future outcomes.
Industry relevance. The evaluated systems come from multiple frontier and open-source families (GPT-5, Grok-4, Claude Sonnet 4, Gemini 2.5 Flash, Llama-4-Scout, OpenAI o3, among 23 total), and the benchmark's modular, searcher-agnostic design lets practitioners swap in retrieval components without changing the evaluation protocol. The finding that even the strongest model does not reach break-even against market prices is a concrete, actionable signal for anyone building forecasting products on top of LLMs.
Future Directions
- Expand and diversify the event pool. The authors state they plan to continue expanding and diversifying Prophet Arena's events in future releases, addressing the current 76% sports skew.
- Fix the three named bottlenecks. Inaccurate event recall, misunderstanding of data sources, and slower information aggregation near resolution are identified as open problems that must be addressed before LLM forecasting becomes reliably superior.
- Advance risk-aware betting and evaluation. The paper mentions a general framework for optimizing betting strategies under different risk preferences (§B.5), a Sharpe-ratio analysis of the betting strategy (§C.2), and a proof about calibration and balanced Yes/No returns (§B.6) — extensions beyond the risk-neutral Average Return used in the main results.
- Probe the remaining mechanistic questions. Section 4 is described as covering knowledge internalization, context construction, and reasoning synthesis, plus robustness and consistency checks on probability elicitation. The truncated content does not report those results, so their conclusions remain open for readers of the full paper.
Target Audience
AI researchers and evaluation scientists studying LLM capabilities, especially those working on forecasting, calibration, or uncertainty quantification; machine learning practitioners building prediction or decision-support systems; and quantitative researchers in finance or market analytics who care about whether model forecasts beat market consensus. Readers with a basic grasp of probability and mean squared error will follow the arguments, since the paper explains the metrics with concrete numerical examples.
Authors’ abstract
Forecasting is not only a fundamental intellectual pursuit but also is of significant importance to societal systems such as finance and economics. With the rapid advances of large language models (LLMs) trained on Internet-scale data, it raises the promise of employing LLMs to forecast real-world future events, an emerging paradigm we call "LLM-as-a-Prophet". This paper systematically investigates such predictive intelligence of LLMs. To this end, we build Prophet Arena, a general evaluation benchmark that continuously collects live forecasting tasks and decomposes each task into distinct pipeline stages, in order to support our controlled and large-scale experimentation. Our comprehensive evaluation reveals that many LLMs already exhibit impressive forecasting capabilities, reflected in, e.g., their small calibration errors, consistent prediction confidence and promising market returns. However, we also uncover key bottlenecks towards achieving superior predictive intelligence via LLM-as-a-Prophet, such as LLMs' inaccurate event recalls, misunderstanding of data sources and slower information aggregation compared to markets when resolution nears.