Skip to content
AI.info

Research

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception

Overview Research area: Natural Language Processing / LLM agents — specifically multi-turn tool (function) calling, temporal reasoning, and human-preference alignment. Technical level: Intermediate (a

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
arXiv
2510.23853
Published
2025-10-27
Authors
Yize Cheng, Arshia Soltani Moakhar, Chenrui Fan, Parsa Hosseini, Kazem Faghih, Zahra Sodagar, Wenxiao Wang, Soheil Feizi

AI summary

Overview

Research area: Natural Language Processing / LLM agents — specifically multi-turn tool (function) calling, temporal reasoning, and human-preference alignment. Technical level: Intermediate (assumes familiarity with LLM agents, tool calling, and preference-optimization fine-tuning, but explains the core problem in plain terms). Scope: This paper defines and measures "temporal blindness" — the failure of LLM agents to account for real-world elapsed time when deciding whether to call a tool — introduces the TicToc dataset to evaluate it, and tests prompting and post-training fixes.

What This Paper Is About

LLM agents that call tools across multiple turns treat the conversation as if no real time passes between messages, even though the outside world keeps changing. This means they sometimes answer from stale context when they should re-check something, and sometimes redundantly re-query facts that could not have changed. The paper measures how poorly current models' tool-call decisions match human judgments about when a tool call is warranted given how much time has elapsed, and tests ways to close that gap.

Key Contributions

  1. Identification of temporal blindness: The paper names and characterizes a previously overlooked failure mode of multi-turn LLM agents — failing to account for elapsed real-world time between messages, leading to either over-reliance on stale context (skipping needed tool calls) or under-reliance (repeating unnecessary tool calls).
  2. The TicToc dataset: A dataset of 1800+ multi-turn user–agent message trajectories (1864 after quality filtering) across 76 scenarios divided into low (29), medium (25), and high (22) time-sensitivity categories, each ending in a user question whose correct response — call a tool versus answer directly — is annotated by humans.
  3. Evaluation and failure analysis of 18 models: A systematic evaluation of proprietary and open-weight LLMs, with and without timestamps, including breakdowns by scenario sensitivity, conversation length, and reasoning behavior, plus an analysis of think–answer mismatches in reasoning traces.
  4. Alignment comparison: A comparison of prompt-based interventions against DPO-based post-training, showing that prompting helps only some advanced reasoning models while targeted post-training produces large alignment gains.

Main Findings

  • Without timestamps, models are near random: In the absence of temporal information, most models' normalized alignment rate (NAR) is similar to random guessing, with the highest NAR reaching just above 55%. The paper notes that a 50% NAR is equivalent to random guessing by definition.
  • Timestamps help only modestly: With timestamp information provided, proprietary OpenAI models and some larger Qwen3 models improve noticeably, but no model achieves an NAR above 65%. In an alternative setting (Appendix B) where the explicit value of the elapsed interval Δt is given instead of an absolute timestamp, little additional alignment improvement is observed.
  • Failures are uniform across time sensitivity: Breaking results down by scenario sensitivity shows models fail across slow and fast-changing environments. Models fail slightly less on medium-sensitivity scenarios on average, but the paper describes this difference as very marginal and the overall alignment rate as low.
  • Tool-call bias varies by model: Without timestamps, most models show higher attempt rates on prefer-Tool samples than on prefer-noTool samples, but each has a distinct bias — Ministral-8B and Llama-3.2-3B tend to invoke tools on nearly all samples, while OpenAI and Qwen models tend to refrain from invoking tools in most cases.
  • Timestamps raise attempt rates on both classes: With timestamps, human-like temporal awareness would mean a higher attempt rate on prefer-Tool samples and a lower one on prefer-noTool samples. Instead, attempt rates rise on both subsets for most models, indicating models struggle to exploit temporal information effectively.
  • Longer conversations trigger more tool calls: Grouping retained samples into short (≤7 turns), medium (8–12 turns), and long (≥13 turns) trajectories, the paper finds a positive correlation between conversation length and tool-call frequency, regardless of whether timestamps are given, with a corresponding dip in NAR on long trajectories. This suggests models use turn count as a heuristic for staleness rather than the explicit time information.
  • Reasoning does not fix the problem: Qwen3 models in reasoning mode show only marginal or no improvement in alignment rate over non-reasoning mode, with or without timestamps.
  • Reasoning traces rarely mention time: Across Qwen3 reasoning models, timestamps appear in fewer than 4% of reasoning traces (ranging from 31 occurrences, 1.03%, for Qwen3-0.6B-Reason to 96, 3.18%, for Qwen3-32B-Reason); the exact keyword "timestamp" appears in less than 1.5% (from 5, 0.17%, to 43, 1.43%); and broader time-related keywords (e.g., "time", "date", "hour") appear in under 15% (270, 8.95%, for 0.6B up to 477, 15.82%, for 8B).
  • Think–answer mismatches drive false positives: Two mismatch types are measured. Type 1 (reasoning decides to call a tool but outputs a direct answer) accounts for 4.11% of false negatives in the 0.6B model and is otherwise negligible. Type 2 (reasoning concludes a direct answer but the final response initiates a tool call) is the major source of false positives — 61.26% of FP errors for Qwen3-1.7B-Reason and approximately 20% for the 0.6B and 4B models (20.00% and 19.75%). For larger models, the mismatch rate falls below 2% and its contribution to FP errors drops under 8%.
  • Prompting has limited reach: A minimal system-prompt reminder ("Note that the environment may be dynamic. Be aware of the time elapsed.") had little to no effect. A stronger prompt with few-shot rules for when tool calls are preferable depending on elapsed time produced a substantial boost for advanced reasoning models such as o3 and o4-mini, but marginal or no effectiveness for most other models.
  • Post-training works: DPO with a dynamic margin, applied to selected open-source models for a single epoch on an approximately 65%:35% scenario-based train/test split, produced what the paper calls massive alignment gains across all trained models.

Methodology in Plain English

The researchers built a scenario taxonomy first: 76 scenarios spanning low-sensitivity environments (regulations, published specifications, archival records — 29 scenarios), medium-sensitivity ones (reservation booking, forecast and condition reports — 25), and high-sensitivity ones (stock markets, competitive bidding, real-time monitoring — 22). Each scenario is either read-only or read+write. Within each setting they defined four trajectory variants capturing distinct follow-up behaviors: for read-only, "Repeated ask," "Comparison," "Retrieve-many, ask-for-one," and "Simple reasoning"; for read+write, "Repeat after failure," "User confirmation," "Repetition of the same request," and "In-context availability / state change."

They hand-wrote one exemplar per variant per scenario, then used GPT-4o to generate candidate trajectories (50 per scenario), which were filtered first by GPT-4.1 acting as an LLM judge and then by human inspection, yielding 1864 high-quality trajectories.

To test time sensitivity, they stamped every message with an ISO 8601 timestamp. Intermediate delays were simulated with sampled per-trajectory pace variables: user reading speed at a mean of 238 words per minute (σ 60), user writing speed at a mean of 3.61 words per minute (σ 0.40, log-normal), and model generation speed at a mean of 40 words per second (σ 16), with tool calls given roughly one-second execution time. The final user question was given three versions per trajectory — Small, Medium, and Large elapsed gaps — with durations scaled to the scenario's sensitivity level (for example, Small spans 3 minutes for all levels; Large spans 3 months for low sensitivity, 3 days for medium, and 3 minutes for high).

Human annotators then judged each of the 1864 × 3 = 5592 samples, choosing among Direct, Lean-Direct, Lean-Tool, or Tool, with scores assigned 0 through 3 respectively. Samples with mean scores between 0.5 and 2.5 were discarded as too uncertain, leaving 3016 retained samples: 1112 prefer-Tool and 1904 prefer-noTool. Inter-annotator agreement was 0.8574 by Krippendorff's alpha.

Evaluation covered 18 proprietary and open-weight LLMs (instruction-tuned versions for Qwen2.5, Ministral, and Llama), run with Temperature = 0 where applicable; Qwen3 reasoning-mode models used Temperature 0.6, TopP 0.95, TopK 20, and MinP 0. Timestamps were injected directly into the chat template for open-weight models (e.g., a user prefix such as <|im_start|>user\n[2025-12-04T10:22:44Z]) and prepended to message text for proprietary models. The headline metric is the Normalized Alignment Rate, the average of true-positive rate and true-negative rate, so that 50% equals random guessing; the paper also reports attempt rate (the proportion of samples where a tool call was attempted). Finally, they tested two mitigation routes: prompt engineering (a minimal reminder and a detailed few-shot rule prompt) and DPO post-training with a dynamic margin on an open-source subset.

Why This Matters

Impact on research: The paper reframes tool-call timing as an alignment problem, not just an accuracy or hallucination problem. It shows that evaluations measuring whether a tool call is correct miss whether it was warranted given elapsed time — and it provides a benchmark, a metric (NAR), and an annotated preference dataset for a question that prior tool-use benchmarks and temporal reasoning benchmarks address only in isolation. It also documents a concrete reasoning failure (think–answer mismatch) specific to this decision context.

Real-world applications:

  • Personal assistants and customer-support agents that must decide whether a cached order status, balance, or booking is still valid or needs re-fetching.
  • Travel and reservation systems where a write action (a booking, a hold) succeeded earlier and the user later asks whether it still holds.
  • Trading, bidding, and real-time monitoring agents where seconds-to-minutes matter and stale reads are costly.
  • Retrieval and research agents that answer from documents or specifications that change slowly, where redundant lookups waste money and latency.

Industry relevance: Every production agent that caches tool output, batches calls, or runs multi-turn sessions implicitly makes the trade-off this paper studies. The finding that turn count — not elapsed time — drives model behavior suggests current systems' "staleness" heuristics are proxies that break in both directions. The result that prompting barely helps most models while DPO post-training produces large gains is a direct argument for spending alignment effort rather than prompt tokens on this problem.

Future Directions

  • Extending to larger models and other modalities: The authors state that their DPO experiments are limited to open-source models with at most 8B parameters due to computational constraints, and that applying targeted DPO to larger-scale models could give further insight. They also flag extending TicToc beyond text-only tool use, for example image retrieval or vision–language tools, as a natural direction.
  • Better evaluation settings for time: The paper tests absolute timestamps and (in Appendix B) explicit Δt values, and finds neither solves alignment. What representation of time would let models actually use it remains open.
  • Understanding and fixing think–answer mismatches: Type 2 mismatches account for the majority of false positives in smaller reasoning models (61.26% for the 1.7B model). Why reasoning conclusions fail to transfer to the emitted action is unresolved.
  • Reducing dependence on prompting: The gap between what detailed prompting achieves on o3 and o4-mini and what it achieves on most other models raises the question of what post-training data or objectives would generalize the effect.

Target Audience

Researchers and engineers working on LLM agents, tool/function calling, and agentic evaluation; alignment practitioners interested in preference data and DPO for behavioral properties beyond harmlessness; and applied teams building multi-turn assistants, booking systems, or monitoring agents where deciding when to call a tool determines both correctness and cost. Readers looking for a first rigorous benchmark of time-aware tool-use decisions and a ready-made annotated dataset and evaluation metric will get the most from it.

Authors’ abstract

Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlooked limitation of these agents is that they, by default, assume a stationary context, failing to account for the real-world time elapsed between messages. We refer to this as "temporal blindness". This limitation hinders decisions about when to invoke tools, leading agents to either over-rely on stale context and skip needed tool calls, or under-rely on it and redundantly repeat tool calls. To study this challenge, we constructed TicToc, a diverse dataset of multi-turn user-agent message trajectories across 76 scenarios, spanning dynamic environments with high, medium, and low time sensitivity. We collected human preferences between "calling a tool" and "directly answering" on each sample, and evaluated how well LLM tool-calling decisions align with human preferences under varying amounts of elapsed time. Our analysis reveals that existing models display poor alignment with human temporal perception, with no model achieving a normalized alignment rate better than 65% when given time stamp information. We also show that naive, prompt-based alignment techniques have limited effectiveness for most models, but specific post-training alignment can be a viable way to align multi-turn LLM tool use with human temporal perception. Our data and findings provide a first step toward understanding and mitigating temporal blindness, offering insights to foster the development of more time-aware and human-aligned agents.

Read the original paper