Skip to content
AI.info

Research

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

Overview Research area: Evaluation methodology for large language model (LLM) agents — specifically, how to allocate a fixed trial budget across interactive agent benchmark scenarios. Technical level:

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets
arXiv
2609.38914
Published
2026-09-30
Authors
Priyanath Maji, Spandan Ghose Chowdhury

AI summary

Overview

Research area: Evaluation methodology for large language model (LLM) agents — specifically, how to allocate a fixed trial budget across interactive agent benchmark scenarios.

Technical level: Intermediate. The paper builds on multi-armed bandit theory (Thompson Sampling) and bandit evaluation practice, but the problem framing and results are accessible to anyone who runs agent benchmarks.

Scope: The paper formulates interactive-agent evaluation as a sequential resource-allocation problem, proposes a risk-aware contextual Thompson Sampling policy, and tests it by offline replay over 70 τ-bench airline scenarios and 824 recorded trials.

What This Paper Is About

Evaluating stochastic LLM agents requires running each scenario many times, which is expensive, yet standard benchmarks spend that budget uniformly — sampling a read-only lookup as often as an irreversible payment action. The authors ask which scenarios should be run and re-run, given a fixed trial budget and unknown failure behavior, so as to find as many high-consequence failures as possible. They separate what is known before execution (a scenario's context features and a fixed impact score) from what is learned during evaluation (observed failure outcomes), and let a bandit policy choose the schedule.

Key Contributions

  1. Formulation. Casts interactive-agent evaluation as a risk-aware sequential allocation problem that separates fixed pre-execution quantities (context vector xᵢ, impact score Iᵢ) from stochastic per-trial observations (failure outcome Fᵢ,ₜ, resource vector Rᵢ,ₜ), with impact-weighted failure discovery as the objective.
  2. Method. Instantiates this as a contextual, risk-aware Thompson Sampling policy and compares it against uniform, random, failure-rate, and static risk-aware baselines plus an oracle upper bound.
  3. Analysis. Uses paired, Holm-corrected significance tests to identify where scenario context and where posterior-based exploration each help, rather than only whether they help.
  4. Cost accounting. Reports discovery per dollar of measured evaluation cost, not only per trial, using the logged user_cost of the GPT-4o simulated user.

Main Findings

  • Small-budget discovery is the headline result. At a budget of 50 trials (6% of the 824-trial corpus), ts-contextual recovers 86.1% of the impact-weighted failure discovery an oracle could achieve, against 24.9% for uniform allocation and 23.9% for random allocation.
  • 3.46x more impact-weighted failures. In raw terms, ts-contextual finds 215.4 versus uniform's 62.2 with the exact same number of trials.
  • All adaptive strategies win at small budgets. At B=50, all four adaptive strategies recover 66–86% of the oracle, supporting hypothesis H1.
  • Context wins the cold start. ts-contextual beats the arm-only variant (identical except it ignores xᵢ) at B ≤ 100, with p_Holm < 0.005; the difference is not distinguishable from noise at larger budgets.
  • Posterior exploration wins at moderate budgets. ts-contextual beats the static risk-aware heuristic by ΔY = +76.1 at B=300 (p_Holm = 4.0 × 10⁻⁹) and +45.8 at B=500 (p_Holm = 8.2 × 10⁻¹⁰), even though it is indistinguishable from ts-arm-only at those budgets — so the advantage comes from posterior-based exploration, not from xᵢ.
  • The advantage shrinks as budget grows. By B=700 (85% of the corpus), all four adaptive strategies converge to 99.7–99.9% of the oracle, while uniform and random remain at 81–82%.
  • 5x discovery per dollar. At B=50, ts-contextual finds 5.0x more impact-weighted failures per dollar than uniform (3207 vs. 636) while spending less in total ($0.067 vs. $0.098). The advantage narrows to 1.16x by B=700.
  • Budget is redirected away from never-failing scenarios. Of the 70 scenarios, 24 (34.3%) never fail in any observed trial. At B=50, ts-contextual spends 2.8% of its budget on these trivial scenarios versus uniform's 34.2%, a 91.8% relative reduction. All four adaptive strategies hold trivial spend below 6% through B=200, but it rises to 21–24% by B=700 as the corpus is exhausted of discoverable failures.
  • The contextual advantage is not redundancy with the impact score. Comparing the full 17-feature vector against a 13-feature variant with the four impact-determining features removed, no budget survives Holm correction (p_Holm ≥ 0.064 everywhere).
  • Hypothesis verdicts. H1 (adaptivity) supported; H2 (context) partially supported and explicitly budget-dependent; H3 (concentration) supported; H4 (cost efficiency) supported at low budgets with diminishing returns at high ones.

Methodology in Plain English

The authors treat each evaluation scenario as a slot machine arm. Pulling an arm means running one trial of that scenario; the payoff is the scenario's fixed impact score multiplied by whether the trial failed. Because failures are rare and consequence varies, the objective is to maximize the total impact-weighted failures found within a fixed number of trials.

Each scenario carries a context vector of 17 features computed statically from the task definition — things like whether it requires a search, a mutation, a payment, a confirmation, or an irreversible action, plus passenger count, insurance sensitivity, complexity, and an ambiguity score. Three of the 17 features are constant in this corpus, leaving 14 active. Each scenario also carries a fixed impact score assigned by a rule: 5 for an expected irreversible action, 4 for mutation plus payment or confirmation, 3 for mutation alone, and 1 for read-only or informational tasks. Impact is never updated from observed outcomes.

Six policies plus an oracle are compared. Uniform cycles through scenarios in a shuffled round-robin. Random picks without regard to history. Failure-rate uses ε-greedy (ε = 0.1) on the observed failure rate. Risk-aware uses the same ε-greedy structure but ranks by impact times observed failure rate. Arm-only Thompson Sampling keeps a Beta posterior per scenario over its failure probability. Contextual Thompson Sampling blends a logistic-regression prediction from the context features, fitted on all pooled observations so far, into those Beta posteriors as pseudo-counts, with a context weight λ = 2 fixed a priori (ablated over λ ∈ {2, 8, 16, 32}).

Rather than executing fresh trials, the authors replay all policies against a single collected corpus of 824 trials from τ-bench's airline domain (70 scenarios), generated with a GPT-4o tool-calling agent at temperature 0.0 against a GPT-4o simulated user. Trial counts per scenario range from 6 to 17. For each of 30 replicates, a scenario's observed trials are shuffled with a replicate-specific seed (seed = 10000s + r) and served without replacement; every policy in a sweep sees the identical shuffled ordering, which is what makes the paired tests valid. Metrics are computed at budgets B ∈ {50, 100, 150, 200, 300, 500, 700} and compared with paired t-tests corrected by the Holm step-down procedure.

Why This Matters

Impact on research. Benchmark protocols almost universally execute a predetermined, uniform trial schedule. This paper makes the schedule itself the optimization target, and reports the first (to the authors' knowledge) decomposition of which mechanism — pre-execution context versus posterior exploration — pays off at which budget. The result that context matters for the cold start and posterior exploration matters mid-budget is a more precise claim than "adaptive allocation works."

Real-world applications:

  • Continuous agent regression testing. Re-evaluating an agent after every model, prompt, or policy change, where each decision can afford only a small slice of the corpus — the exact regime the paper targets.
  • Safety-critical tool use. Prioritizing evaluation toward scenarios involving irreversible actions, payments, or policy-sensitive operations so higher-stakes failures surface earlier.
  • Cost-constrained evaluation pipelines. Teams whose evaluation spend is billed per trial against live tool-calling environments and LLM-based user simulators can extract more discovery per dollar.
  • Benchmark design. Deciding how many tasks and trials a benchmark needs for a stable reliability decision, given that adaptive allocation dominates uniform at every budget short of exhausting the corpus.

Industry relevance. Evaluation cost scales directly with repeated trials, and the paper shows that uniform allocation spends roughly a third of its budget on scenarios that never fail in any observed trial. The reported 5x improvement in discovery per dollar at the 50-trial budget, at lower absolute spend, is a practical argument for changing how evaluation budgets are scheduled rather than how agents are built.

Future Directions

  1. Learned, trajectory-conditioned impact models. Replacing the rule-based Iᵢ, which cannot infer the severity of a new failure from its trajectory, with a learned impact model.
  2. Hierarchical or fully Bayesian generalized linear bandits. Modeling each scenario as an independent arm ignores correlations; scenarios sharing a policy or operation type likely share failure modes. The current contextual variant uses a plug-in empirical-Bayes approximation rather than a full posterior over the logistic link, a simplification more likely to matter at larger corpus sizes.
  3. Cost-adjusted selection. No policy in this study normalizes its score by estimated cost, so the observed cost efficiency is a byproduct of failure-aware selection rather than an optimized objective.
  4. Broader objectives and replication. Extending the objective toward failure diversity and calibrated reliability estimation, and replicating across other benchmarks — customer support, healthcare, finance, or software engineering — since results here come from a single airline environment with 70 scenarios.

Target Audience

Researchers and practitioners who run or design evaluations for interactive LLM agents, especially those working under tight trial budgets. It is also relevant to applied machine learning engineers building continuous evaluation pipelines and to benchmark designers deciding how many tasks and trials a reliability measurement needs. Readers without a bandit background can follow the problem framing and results, though the method section assumes comfort with Beta posteriors and Thompson Sampling; readers seeking the full statistical detail will want the paired-test tables and the appendices on feature composition, the trivial-scenario sweep, per-scenario allocation, and the context-weight ablation.

Authors’ abstract

Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We instead formulate evaluation as a sequential allocation problem. Given a fixed trial budget and a set of scenarios whose failure behavior is unknown, which scenarios should be run, and run again? We propose a risk-aware contextual Thompson Sampling policy that combines a pre-execution scenario context vector and a fixed impact score with the failure outcomes observed during evaluation, and we test it by offline replay over 70 $τ$-bench airline scenarios and 824 recorded trials. Our main result is at the smallest budget: with only 50 trials ($6\%$ of the corpus), the policy recovers $86\%$ of the impact-weighted failures an oracle could find, compared to $25\%$ for uniform allocation. It discovers $3.5\times$ more impact-weighted failures (215.4 vs. 62.2) with the same number of trials, delivers $5\times$ the discovery per dollar, and cuts the budget wasted on scenarios that never fail from $34\%$ to $2.8\%$. The rest of our analysis demonstrates and qualifies this result: a budget sweep shows the advantage shrinks as the budget approaches the corpus size, and paired significance tests show that scenario context helps mainly at small budgets while posterior-based exploration helps at moderate ones. Risk-aware adaptive allocation therefore helps most exactly where evaluation budget is scarcest.

Read the original paper