Skip to content
AI.info

Research

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Overview Research area: LLM agents, human–AI interaction, agent evaluation, mixed-initiative and proactive assistance. Technical level: Advanced (formal agent definitions, a design-space taxonomy, an

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
arXiv
2609.37267
Published
2026-09-29
Authors
Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang

AI summary

Overview

  • Research area: LLM agents, human–AI interaction, agent evaluation, mixed-initiative and proactive assistance.
  • Technical level: Advanced (formal agent definitions, a design-space taxonomy, an evaluation testbed, LLM-as-judge metrics, and a human study).
  • Scope: The paper proposes three joint design principles (3T) for proactive LLM agents, maps them onto a five-dimension design space and the system components needed to realize them, and introduces Proactivity-Gym, a multi-day simulation testbed with evaluations across 23 model–harness configurations plus a 30-participant human study.

What This Paper Is About

Proactive LLM agents are meant to turn idle compute into useful support before the user asks, but correct work can still be badly timed, misaligned with what the user wants, or damaging to trust. The authors argue that current research and deployment remain largely reactive and measure success only by task fulfillment, so failures in timing and trust get hidden. The paper's goal is to define what proactivity requires, lay out the design and system choices it implies, and build an evaluation testbed that measures task capability, compute allocation, and user trust together.

Key Contributions

  1. Principles (3T): A blueprint for proactive LLM agents built on three joint objectives — Task Capability (TC), Temporal Allocation (TA), and Trust (TR) — which the authors argue are commonly conflated or overlooked, with a weighted joint design objective combining scores for each.
  2. Technical layers: A formalization of the proactivity design space along five dimensions (task scope, anticipation horizon, activation trigger, processing timing, intervention depth), together with the situation modeling and system modeling (backbone LLM plus harness) required to realize it.
  3. Proactivity-Gym: A simulation-based evaluation testbed with 10 hand-crafted multi-day scenarios (each 7–10 simulated days) and three user personas per scenario, using a simulated clock, stateful tool environments, and persona-conditioned simulated users, evaluated with TC, TA, TR-D, and TR-J metrics.
  4. Empirical evidence: Evaluations across 23 model–harness configurations showing substantial deficits in current agents, plus a 30-participant human study indicating that humans and LLM judges weigh intervention-depth misalignment very differently.

Main Findings

  • Strongest agents still fall short: Claude Opus 5 achieves the highest average TC score of 65.1, but reaches only 51.7% on TA and 52.4% on TR-D. All other models remain below 20% on TA, with TR-D ranging from 39% to 51%. Agents particularly struggle to defer competing work.
  • Model size helps within families: Within model families, larger models score higher on all four metrics.
  • Harness effects depend on the model and objective: Across the five open-weight models, switching from Claude Code to OpenClaw lowers TC by 3.6 points on average, with changes ranging from −11.1 for Gemma 4 12B to +0.5 for Qwen 3.5 27B. For Claude Opus 5, the same switch raises TC by 4.6% and TA by 21.1%, while lowering TR-D by 3.8% and TR-J by 0.20.
  • TC gains do not uniformly improve TA or TR: TC correlates weakly with TA (r = 0.31) and TR-D (r = 0.34), but more strongly with TR-J (r = 0.70).
  • More reasoning does not fix trust: Raising GPT 5.6-Sol's reasoning effort from none to xhigh raises TC from 50.0 to 58.9, while both TR metrics plateau.
  • LLM judges overlook intervention misalignment: Frontier models receive relatively high TR-J scores, especially for understandability and perceived technical competence; Claude Opus 5 scores 4.72 and 4.60 on those two dimensions while scoring 52.4% for TR-D. This suggests LLM-as-judge scores are tolerant of intervention-depth misalignment and weight TC more heavily when assessing trust.
  • Humans value correctness, but also alignment: Participants preferred correct actions or responses (TC↑) in 92.2% of comparisons, and when content quality was held constant they selected agents aligned with the user's intervention preference (TR↑) in 88.3% of comparisons.
  • Timing can rescue otherwise-declined assistance: Willingness to accept proactive assistance increased from 26.7% for correct assistance competing with ongoing tasks to 97.8% for sleep-time assistance, even when the output required revision the next morning. Even when immediate assistance required only 10 minutes of review with no resource or deadline conflict, 64.4% preferred sleep-time assistance.
  • Trust is easier to lose than rebuild: An aligned-to-misaligned (A→M) transition is followed by a 1.86-point decrease on average, compared with a 1.27-point increase following the reverse transition. In the A-M-A sequence, mean trust recovers only to 3.43, below its initial level of 4.43, even though participants rate the final intervention itself as appropriate (4.71).

Methodology in Plain English

The authors start from a definition of proactivity as anticipatory behavior that addresses gaps in a user's needs, opportunities, or problems without an explicit request. They split what a proactive agent must get right into three objectives: doing useful work, spending compute at the right time, and maintaining the user's confidence and willingness to rely on it. They then organize the practical choices into five dimensions — what work to pursue, when to start it, what triggers it, when to process it, and how far to proceed autonomously — with three intervention levels (prepare, suggest, execute) adapted from the Interface-Proactivity continuum. To show what a system needs, they describe situation modeling (representing the user, their goals, attention, trust, and the environment's tasks, deadlines, and resources) and system modeling (the backbone LLM and a harness that tracks pending work, allocates compute, and controls permissions), formalized as a loop where a representation is built from context, an action is selected, and context is updated with new observations.

For evaluation, they build Proactivity-Gym: scenarios with 7–10 simulated days each, timestamped events and tasks, stateful tools, three personas per scenario that differ in preferred intervention depth, and a simulated clock. Scenarios include latent needs inferrable from prior interactions, competing tasks under resource, availability, and deadline constraints, and NOOP cases where no intervention is needed. They test nine models (GPT 5.6-Luna, GPT 5.6-Sol, Claude Sonnet 5, Claude Opus 5, Qwen 3.5-2B/9B/27B, Gemma 4-12B/31B) across three harnesses (OpenClaw, Claude Code, Codex), yielding 23 model–harness combinations, with three runs per scenario-persona pair and reasoning disabled in the main experiments. TC (0–100) combines rule-based and LLM-as-judge assessments; TA (%) measures whether urgent work is prioritized and competing tasks deferred; TR-D (%) measures agreement between the agent's intervention depth and the user's preference; TR-J (1–5) averages ratings of five trust constructs (understandability, technical competence, reliability, personal attachment, faith) based on persona and cumulative interaction history. Judge scores are averaged over two independent judges, Qwen 3.8-27B and Gemini 3.8-Flash. A separate study with 30 students and working professionals reviewed 14 scenarios, mostly adapted from Proactivity-Gym including four week-long interaction logs, using paired comparisons and five-point rating scales, with trust examined at Days 2, 4, and 7.

Why This Matters

  • Impact on research: The paper argues that evaluating proactive agents on isolated responses or static settings hides the consequences that matter, and that existing benchmarks mostly cover task capability while omitting compute allocation and trust. It supplies shared vocabulary (3T), a design space, and a testbed that lets failures in timing and trust be diagnosed separately from task failures, and it documents a gap between LLM-judge trust scores and human trust judgments.
  • Real-world applications:
    • Personal assistants that prepare work overnight and present results at the next interaction, rather than interrupting during focused periods.
    • Research and coding workflows where background jobs (for example, running missing ablations while GPUs are free) are scheduled around the user's own compute use.
    • Everyday planning agents that monitor external events such as flight delays and rebook or update plans after approval.
    • Scheduling and reminder agents that must respect an explicit standing request for confirmation, as in the paper's language-class example (a 15-minute review session with a reminder 10 minutes beforehand).
  • Industry relevance: Harness choice changes outcomes in ways that are model- and objective-dependent — a switch that raises TC or TA can lower TR-D and TR-J — so teams building agent products cannot assume better task scores mean better user experience. The finding that users accept sleep-time assistance even when output needs correction (97.8% versus 26.7%) suggests compute scheduling is a product lever, not just an infrastructure detail, and the finding that one misaligned intervention sharply reduces trust implies deployment permissions deserve deliberate control.

Future Directions

  • Joint 3T optimization in agent design: The paper notes that current systems largely optimize task capability and states that the joint formulation remains to be realized in practice.
  • Extending the gym: The authors describe their testbed as initial, with simplified resource constraints and user models, and discuss possible extensions in their appendix.
  • Longer-term real-world evaluation: Moving beyond simulation to real deployment is listed as future research.
  • Reconciling LLM judges with human perception: Since LLM-as-judge trust scores are relatively insensitive to intervention-depth misalignment while humans show sharp trust declines, better trust evaluation methods are needed.
  • Data and code release: The authors state they plan to release code and Proactivity-Gym data upon publication.

Target Audience

Researchers and practitioners working on LLM agents, agent harnesses, and personal assistants who need to evaluate behavior beyond single-turn task success; HCI and trust-in-automation researchers interested in mixed-initiative and proactive systems; and engineering teams building or deploying always-available agents who must decide when to prepare, suggest, or act autonomously.

Authors’ abstract

Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.

Read the original paper