Skip to content
AI.info

Research

STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls

Overview Research area: Design-time decision frameworks for AI system architecture, specifically choosing between direct LLM calls, guided AI assistants, and fully autonomous agentic AI in enterprise

STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
arXiv
2512.02228
Published
2025-12-01
Authors
Shubhi Asthana, Bing Zhang, Chad DeLuca, Ruchi Mahindru, Hima Patel

AI summary

Overview

  • Research area: Design-time decision frameworks for AI system architecture, specifically choosing between direct LLM calls, guided AI assistants, and fully autonomous agentic AI in enterprise settings.
  • Technical level: Intermediate. The framework uses DAG-based task decomposition and explicit scoring equations, but the paper explains each component and its inputs in plain terms.
  • Scope in one sentence: The paper introduces STRIDE (Systematic Task Reasoning Intelligence Deployment Evaluator), a five-stage framework that scores a task description and recommends whether it needs a stateless LLM call, a guided AI assistant, or full agentic autonomy, evaluated on 30 real-world tasks.

What This Paper Is About

As AI systems move from stateless LLM calls to autonomous, goal-driven agents, teams tend to deploy agentic AI even when a simpler modality would work, which raises cost, complexity, and risk. The paper addresses the missing design-time question of when agents are truly necessary, rather than how well they perform after deployment. STRIDE answers that question by decomposing a task, scoring its reasoning, tool, state, and risk needs, attributing its dynamism, and checking whether self-reflection is required before issuing a modality recommendation.

Key Contributions

  1. Introduction of STRIDE, described as the first design-time framework for AI modality selection, shifting the decision from deployment time to the design phase (a "shift-left" tool).
  2. A quantitative Agentic Suitability Score (ASS) computed per subtask with dynamism attribution, designed to balance autonomy benefits against cost and risk.
  3. Evaluation on 30 real-world tasks across SRE, compliance, and enterprise automation (the experiments section also lists customer support), reporting reduced agentic over-deployment by 45% while improving expert alignment by 27%.
  4. A five-dimension analysis pipeline combining structured task decomposition, dynamic reasoning and tool-interaction scoring, dynamism attribution via a True Dynamism Score, self-reflection requirement assessment, and agentic suitability inference with persona-aware recommendations.

Main Findings

  • Overall accuracy and efficiency: Across all 30 tasks, STRIDE achieved 92% accuracy, reduced unnecessary agent deployments by 45%, and delivered 37% lower compute/API usage compared to always deploying agents. Table 2 lists the precise figures as 92.0% accuracy, 45.3% over-engineering reduction, and 37.1% resource savings.
  • Baseline comparison: The Naive Agent baseline (always deploy agentic AI) scored 33.3% accuracy with 0% over-engineering reduction and 0% resource savings. The Heuristic Threshold baseline (deploy agents only when reasoning depth ≥ 2 and tool requirements ≥ 2) scored 68.0% accuracy, 27.5% over-engineering reduction, and 18.2% resource savings. STRIDE outperformed both.
  • Domain-wise accuracy: STRIDE reached 95% accuracy in SRE, 91% in compliance, 89% in automation, and 93% in customer support, which the authors interpret as generalization across heterogeneous tasks without overfitting to a single domain.
  • Expert validation: Experts fully agreed with STRIDE in 78% of cases, partially agreed in 15%, and disagreed in 7%, producing a 27% improvement in expert alignment over the Heuristic Threshold baseline.
  • Extended human collaboration: For SRE, three Kubernetes incident response experts engaged iteratively over a six-month period (March–August 2025). For compliance, two legal verification experts participated over 1–2 months (May–June 2025).
  • Ablation results: Full STRIDE scored 92.0% accuracy, 45.3% over-engineering reduction, and 37.1% resource savings. Removing task decomposition dropped accuracy to 83.0%; removing the True Dynamism Score dropped it to 80.0%; removing TDS weighting gave 81.3%; removing self-reflection reduced accuracy to 76.0%; removing human-in-the-loop gave 85.7%.
  • Component importance: Removing task decomposition reduced accuracy by 9%, removing the True Dynamism Score reduced accuracy by 12%, and removing self-reflection produced the largest drop, with accuracy falling to 76%.
  • Error direction: Errors arose mainly in borderline scenarios such as multi-document summarization, where dynamism was underestimated. STRIDE sometimes recommended assistants when experts preferred agents, but never the reverse, avoiding costly over-engineering.
  • Representative task scores: A currency lookup received a True Dynamism Score of 0.10 and was routed to LLM_CALL; meeting summarization received 0.35 and was routed to AI_ASSISTANT; travel itinerary planning received 0.78, Kubernetes incident analysis 0.85, and legal compliance verification 0.80, all routed to AGENTIC_AI.
  • Historical pattern example: "Search Flights" appears as the starting point in 85% of travel planning tasks in the system's stored patterns.
  • Weighting examples: The framework uses domain-adapted weights such as 0.4 for reasoning-heavy tasks (itinerary planning), 0.3 for tool-intensive workflows, 0.2 for context-dependent operations, and 0.1 for risk-sensitive applications like compliance.
  • Acknowledged limitation: STRIDE's scoring functions are heuristic by design, striking a balance between interpretability and generality.

Methodology in Plain English

STRIDE takes a plain task description and runs it through five stages.

First, it breaks the task into subtasks using a fine-tuned LLM with specialized prompting. It looks for action verbs (search, validate, analyze) and target nouns (flights, budget, data), then links subtasks into a directed acyclic graph based on temporal ordering, data flow, and semantic role labeling. For "Plan a 5-day travel itinerary," this yields subtasks such as Search Flights, Find Hotels, Budget Planning, and Activity Research.

Second, each subtask is scored on four dimensions: reasoning depth (none, medium, or deep), tool need (none, single, or multiple), state or memory requirement (none, ephemeral, or persistent), and risk (compliance violations, computational overhead, infinite loop potential). These combine into an Agentic Suitability Score using weighted terms. The weights are tuned by grid search on labeled historical task data, refined through reinforcement learning from deployment outcomes, and calibrated further by expert feedback.

Third, STRIDE attributes the source of variability in a task. Model-induced variability (prompt ambiguity, stochastic randomness) can usually be fixed with prompt engineering or temperature control. Tool-induced variability (API volatility, changing response formats) calls for error handling and retries. Workflow-induced variability (conditional branching, changing environmental conditions) is the category that genuinely indicates agentic need. A True Dynamism Score combines workflow variability and tool volatility while subtracting model instability, so that a high score means adaptivity is truly required.

Fourth, STRIDE checks whether self-reflection is needed, which happens when a task has mid-execution decision points or relies on nondeterministic tools requiring validation. A decision rule triggers reflection hooks such as error recovery, re-planning, or ReAct-style loops when the dynamism score exceeds a threshold and conditional branches, nondeterministic tools, or mid-execution validation are present.

Fifth, the subtask features are aggregated into a task profile and passed to a classifier that queries a knowledge base of historical patterns, producing the final modality recommendation along with persona-tailored justification (tool configurations for developers, architectural summaries for managers).

Why This Matters

Impact on research. Existing benchmarks such as AgentBench, ITBench, ToolBench, SWE-Bench, Gorilla, HuggingGPT, and ReAct measure how well agents perform after deployment. STRIDE addresses the complementary question of whether agents are needed at all before deployment, reframing agent adoption from intuition-driven to a structured, repeatable decision process. The authors argue that just as scaling laws guided model development, an analogous structured perspective is needed for environmental and task scaling.

Real-world applications.

  • Site reliability engineering: Kubernetes incident analysis, where change events must be correlated with active alerts to find root causes, scored a True Dynamism Score of 0.85 and was routed to agentic AI.
  • Compliance and legal verification: Evaluating documents for non-compliant sections and suggesting corrections, scored at 0.80, where experts noted that assistants often fail to capture regulatory edge cases.
  • Enterprise automation: Structured processes that need guidance and human oversight but not autonomous decision-making.
  • Customer support: One of the four domains in the robustness evaluation, reaching 93% accuracy.

Industry relevance. The framework is positioned as a "shift-left" decision tool for enterprise AI workflows, providing defensible criteria for balancing capability, efficiency, computational cost, and risk. The authors frame it as a guardrail for responsible AI deployment: by preventing over-engineering, it reduces unnecessary surface area for errors, governance failures, security exposure from uncontrolled tool use, and hidden costs. The paper also distinguishes three risk categories motivating the work: overengineering, security and compliance risks from uncontrolled tool and API use, and system instability from recursive loops and unbounded workflows.

Future Directions

  • Multimodal expansion: Extend evaluation beyond the 30 tasks to include multimodal tasks involving vision and audio.
  • Reinforcement learning for weight tuning: Integrate RL more directly into tuning the scoring weights.
  • Enterprise-scale validation: Validate STRIDE at enterprise scale rather than on the current 30-task set.
  • Integration with existing benchmarks: Use STRIDE as a design-time filter that guides which tasks should be benchmarked with agents, or as a planning tool embedded into enterprise AI workflows.
  • Open question on scoring rigor: The authors acknowledge that the scoring functions are heuristic by design, leaving room to strengthen the formal basis while preserving interpretability.

Target Audience

Enterprise architects and engineering leaders deciding how much autonomy to build into AI systems; AI product managers and platform teams choosing between LLM calls, assistants, and agents; SRE and compliance practitioners evaluating AI for high-stakes workflows; and researchers working on agent evaluation, task decomposition, and responsible AI deployment who want a design-time complement to post-deployment benchmarks.

Authors’ abstract

The rapid shift from stateless large language models (LLMs) to autonomous, goal-driven agents raises a central question: When is agentic AI truly necessary? While agents enable multi-step reasoning, persistent memory, and tool orchestration, deploying them indiscriminately leads to higher cost, complexity, and risk. We present STRIDE (Systematic Task Reasoning Intelligence Deployment Evaluator), a framework that provides principled recommendations for selecting between three modalities: (i) direct LLM calls, (ii) guided AI assistants, and (iii) fully autonomous agentic AI. STRIDE integrates structured task decomposition, dynamism attribution, and self-reflection requirement analysis to produce an Agentic Suitability Score, ensuring that full agentic autonomy is reserved for tasks with inherent dynamism or evolving context. Evaluated across 30 real-world tasks spanning SRE, compliance, and enterprise automation, STRIDE achieved 92% accuracy in modality selection, reduced unnecessary agent deployments by 45%, and cut resource costs by 37%. Expert validation over six months in SRE and compliance domains confirmed its practical utility, with domain specialists agreeing that STRIDE effectively distinguishes between tasks requiring simple LLM calls, guided assistants, or full agentic autonomy. This work reframes agent adoption as a necessity-driven design decision, ensuring autonomy is applied only when its benefits justify the costs.

Read the original paper