Research
Do Large Language Models (LLMs) Understand Chronology?
Do Large Language Models (LLMs) Understand Chronology? Overview Research area: Chronological and temporal reasoning in large language models, motivated by the use of LLMs as forecasting tools in finan

- arXiv
- 2511.14214
- Published
- 2025-11-18
- Authors
- Pattaraphon Kenny Wongchamcharoen, Paul Glasserman
AI summary
Do Large Language Models (LLMs) Understand Chronology?Overview
Research area: Chronological and temporal reasoning in large language models, motivated by the use of LLMs as forecasting tools in finance and economics.
Technical level: Intermediate. The paper is readable without deep NLP background, but it assumes familiarity with LLM prompting, reasoning models, and rank-correlation metrics (Spearman's ρ, Kendall's τ, normalized Cayley distance, exact match rate).
Scope in one sentence: The authors build a knowledge-verified testbed of three escalating chronological tasks — ordering, conditional sorting, and anachronism detection — over historical facts the models already know, and evaluate GPT-4.1, Claude-3.7 Sonnet (with and without Extended Thinking), and GPT-5 across multiple reasoning-effort settings.
What This Paper Is About
Methods for avoiding look-ahead bias in LLM-based financial forecasting often rely on prompts such as "use only information from before 2016," which implicitly assumes the model understands what chronological ordering means. This paper steps back and tests that assumption directly: it asks whether LLMs can correctly sequence, filter, and detect anachronisms in events drawn from periods the models were trained on. The goal is to measure chronological understanding of already-known facts "in the wild," rather than to isolate a specific component of temporal reasoning or to measure data leakage itself.
Key Contributions
-
A knowledge-verified chronology testbed. Before any ordering trial, every candidate item is screened by a separate query asking for its year; an item is used only if the model's answer exactly matches the canonical year. This isolates chronological reasoning skill from gaps in factual recall. The testbed covers three task families of increasing difficulty: basic chronological sorting, conditional sorting (filter then order), and anachronism detection.
-
A documented dissociation between strict accuracy and rank quality. The paper shows that exact match ordering collapses as lists grow while rank correlations (Spearman's ρ, Kendall's τ) remain high, indicating that LLMs preserve local order but fail to maintain a single globally consistent timeline.
-
Evidence that conditional-sorting failures come mostly from the filtering step, not the ordering step, which means reporting rank correlations on the valid subset alone can overstate real performance.
-
Demonstration that explicit deliberation flips the pattern. Claude 3.7 Sonnet with Extended Thinking and GPT-5 at medium/high reasoning effort achieve flawless (100%) exact-match ordering at every tested list length and perfect conditional sorting, while low/minimal reasoning effort degrades with longer lists, mirroring the non-reasoning models.
Main Findings
-
Exact match collapses with list length, while rank correlation stays high (20th-century events). On the filtered corpus of N = 100 events (one per calendar year of the twentieth century, 20 trials per list size), exact match rate is 1.00 at n = 2, about half the time at n = 5 (EMR ≈ 0.45), and 0.00 at n = 20, 50, and 100. Spearman's ρ and Kendall's τ stay relatively flat from 5 to 50 items but drop at 100 (ρ ≈ 0.786, τ ≈ 0.661). Normalized Cayley distance rises monotonically from 0.000 at n = 2 to 0.909 at n = 100, meaning more swaps are needed to repair longer permutations.
-
GPT-4.1 beats random guessing but not on strict matching. Against 1,000 uniformly random permutations per list length, GPT-4.1 ranks in the 95th–100th percentile for rank correlation and normalized Cayley distance. Exact match converges to the 50th percentile once n ≥ 20, because random permutations almost never achieve a perfect match either.
-
Errors concentrate in the middle of lists. Across ground-truth positions, mean absolute rank difference (MARD) stays below two positions for all slots at n = 2, 5, 10, but for n = 20 and 50 errors grow to 3–6 positions on average and vary sharply with position, with wider variability bands.
-
Wide temporal gaps help but do not fix exact match. In the wide-gap variant (single-year events spanning Years 1–2025, with target gaps Δtgt ∈ {50, 100, 200} years, list sizes n ∈ {2, 5, 10, 15, 20, 24}), rank correlations stay high for n ∈ {10, 15, 20, 24} (ρ ≈ 0.96–0.97, τ ≈ 0.89–0.92). Exact match still falls from 1.00 (n = 2) to 0.55 (n = 5) to 0.20 (n = 10) and 0.00 for n ≥ 15. Normalized Cayley distance rises from 0.125 (n = 5) to 0.411 (n = 24), i.e. roughly 0.41 × (n − 1) swaps on average.
-
Omissions and hallucinations appear on U.S. presidents. Of 200 trials, 42 omit at least one name from the prompt, first appearing at n = 15 (3/20 trials), peaking at n = 25 and n = 30 where more than half the trials omit a name. 64 trials return names not present in the prompt; hallucinations affect 85% of n = 40 trials and 95% of the full n = 43 trials, but there are no omissions or hallucinations for n ≤ 10. Early figures such as James K. Polk and Andrew Jackson are most often missing; Obama, Biden, and both Bushes are frequently added.
-
U-shaped correlation pattern on presidents. After cleaning, exact match for GPT-4.1 falls from 0.96 (n = 2) to 0.00 at n = 30, 35, 40, and 43, while Spearman's ρ and Kendall's τ dip as lists lengthen and then rebound (ρ = 0.992/τ = 0.986 at n = 35; ρ = 1.000/τ = 1.000 at n = 43). The authors interpret this as the model "knowing" the complete presidential timeline better than arbitrary subsets of it. MARD per position stays modest for r ≤ 20 and peaks around r ≈ 30–35; accuracy is highest for 18th-century presidents, degrades through the 19th and 20th centuries, and improves again for the 21st-century cohort.
-
Explicit deliberation produces perfect ordering. Claude 3.7 Sonnet with Extended Thinking achieves exact match = 1.00 at every list size, with normalized Cayley distance indistinguishable from zero, and suppresses missing and hallucinated names entirely. Without Extended Thinking, Claude 3.7 does not reliably outperform GPT-4.1 and actually lags it at n = 5.
-
GPT-5 shows a sharp reasoning-effort threshold. At medium and high reasoning effort, GPT-5 achieves perfect chronological ordering at all list sizes (exact match = 1.00, ρ = τ = 1.00). The low setting is near-perfect with a single dip at n = 25 (exact match = 0.95). Minimal and no-reasoning GPT-5 variants mirror non-reasoning models: ρ and τ ≥ 0.94 but exact match collapses as lists grow.
-
Conditional sorting fails mainly at the filter. With GPT-4.1, known attributes such as Ohio (7 names), Virginia (8 names), Massachusetts (4 names), and EvenYears (22 names) almost never produced perfect filtering (single-digit success rates over 100 trials). The combined OhioOrVirginia
Authors’ abstract
Large language models (LLMs) are increasingly used in finance and economics, where prompt-based attempts against look-ahead bias implicitly assume that models understand chronology. We test this fundamental question with a series of chronological ordering tasks with increasing complexities over facts the model already knows from pre-training. Our tasks cover (1) chronological ordering, (2) conditional sorting (filter, then order), and (3) anachronism detection. We evaluate GPT-4.1, Claude-3.7 Sonnet, with and without Extended Thinking (ET), and GPT-5 across multiple reasoning-effort settings. Across models, Exact match rate drops sharply as sequences lengthen even while rank correlations stay high as LLMs largely preserve local order but struggle to maintain a single globally consistent timeline. In conditional sorting, most failures stem from the filtering step rather than the ordering step, but GPT-5 and Claude-3.7 Sonnet with Extended Thinking outshine normal models significantly. Lastly, anachronism detection is found to be the easiest task for the LLMs but performance still declines with increasingly overlapping timelines or entities. Overall, our main contribution is showing that allocating explicit reasoning budget helps with chronological ordering with GPT-5 at medium/high reasoning effort achieving flawless ordering at all lengths and perfect conditional sorting (both self-filtered and given-subset), whereas low/minimal effort degrades with longer lists, mirroring earlier models. Our findings delineate limits of current LLMs on chronological tasks, providing insights into task complexity, and demonstrate scenarios in which reasoning helps. These patterns are important for the real-time application of LLMs in finance. We release all code and evaluation templates to support full reproducibility.