Skip to content
AI.info

Research

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

Overview Research area: Evaluation methodology for large language model (LLM) agent marketplaces (agent-to-agent commerce), sitting at the intersection of causal inference, construct validity, and mul

arXiv
2609.01519
Published
2026-09-01
Authors
Peiying Zhu, Sidi Chang

AI summary

Overview

Research area: Evaluation methodology for large language model (LLM) agent marketplaces (agent-to-agent commerce), sitting at the intersection of causal inference, construct validity, and multi-agent simulation.

Technical level: Intermediate. The economic framing (buyer surplus, seller profit, welfare) and the statistical machinery (bootstrap intervals, variance decomposition, exact sign-flip tests) are explained plainly enough for a general reader with some familiarity with LLM agents, but the paper assumes comfort with experimental design language.

Scope: A forensic audit of a single multi-turn buyer–seller hotel-negotiation testbed, in which the authors re-run their own previously positive guardrail result and show that the apparent effect is not identified by the original protocol.

What This Paper Is About

Interactive LLM market simulations produce outputs that look economic — prices, profits, consumer surplus, welfare — even when the simulation does not actually instantiate the behavior its authors claim to measure. The paper asks what must be checked before a welfare claim from such a simulation can be believed. Using their own multi-turn hotel buyer–seller testbed as a case study, the authors tear down an initial positive result, propose a four-part validity contract, and show their own headline number is either Invalid or Inconclusive under that contract.

Key Contributions

  1. A documented scaffold sensitivity. An apparent guardrail effect changes sharply — including a sign reversal at the 3B model — once guarded and unguarded cells share one offer schema and one buyer choice rule.

  2. A quantified single-generation instability result, without pseudo-replication. Replications are averaged within profile-condition, and profiles (not generation rows) are the paired unit. An exact profile-level sign-flip placebo finds the selected effect compatible with label exchangeability (p = 0.50).

  3. An exposed incentive-validity gap. Prompt-defined seller roles do not respond monotonically to a stronger profit instruction, and scripted positive controls show that the welfare effect of guardrails changes sign depending on the assumed seller technology.

  4. A compact evaluation contract with a three-way decision rule. Four predicates (incentive validity, protocol isolation, stochastic stability, accounting completeness) gate substantive interpretation, returning Invalid for a known construct violation and Inconclusive when precision or coverage is insufficient. A reusable nine-item checklist accompanies the artifact.

Main Findings

  • The original headline was large and uniform. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B–14B ladder.

  • The original implementation changed more than the guardrails. Guarded and unguarded agents were given different offer schemas and different choice procedures, so policy and transaction-construction scaffold varied together.

  • Fixing the protocol collapses the effect. Holding schema and buyer chooser fixed changes the same paired contrasts to +7.2 (95% CI [-8.1, 23.8]), −13.9 ([-26.2, −5.2]), and +23.8 ([-1.5, 56.6]). The estimate shrinks by 92% at 1.5B and reverses sign at 3B; at 14B the point estimate stays positive but is not resolved from zero.

  • The surviving 14B positive contrast is an isolated super-additivity. The 2×2 interaction (Both − Info − Conduct + None) is −2.5 at 1.5B and −0.3 at 3B, but +35.8 at 14B. At 14B both single-rule main effects are negative (−3.5 and −8.4), so an additive model predicts welfare of 93.5, while observed Both welfare is 129.3. This appears at only one model size with one generation per profile-condition.

  • The 14B channel runs mainly through completion. Both raises completion from .77 to .90, while accepted-trade welfare rises only from 137.5 to 143.6. Holding accepted-trade welfare at the None level, the completion change accounts for 18.3 of the 23.8 aggregate-welfare units, or 77%.

  • Selected profiles overstate effects. The four largest 14B single-generation effects were +251, +191, +270, and +204, averaging +229. With three new generations per cell, the profile effects become +83.7, −68.3, 0, and +135, averaging +37.6 with a profile-bootstrap interval of [-34.2, 109.3].

  • Low temperature did not make one rollout stable. Mean within-cell generation SD is 47.8, the approximate 80%-power minimum detectable paired effect is 180.9 (nearly five times the repeated point estimate), and the exact profile-level sign-flip test gives a two-sided p = 0.50 (one-sided p = 0.25).

  • Generation noise dominates the variance decomposition. In the balanced profile × condition × replicate table, generation residuals account for 49.9% of sum-of-squares variation, profile effects 24.1%, profile-by-condition interaction 21.1%, and condition 4.9%. These percentages apply only to the selected four-profile probe.

  • The assumed strategic seller is not validated. The profit-pressure prompt increases rent-extraction questions relative to the standard prompt (0.83 vs. 0.17 per dialogue) but produces much less seller profit (33.8 vs. 57.5) and a lower profit share (0.37 vs. 0.56). The compliance prompt yields the least profit, but the intended middle-to-high ordering fails. With only 3 profiles, C1 is judged Inconclusive rather than a known violation.

  • Scripted positive controls locate the missing mechanism. A profit-maximizing seller extracts all buyer surplus in None while selecting a first-best bundle, yielding mean welfare of 169.1 and 100% of first best. Adding both guardrails transfers 44.0 to buyers, removes 68.5 from sellers, and reduces welfare by 24.5 (−15.4 percentage points of first best). An inefficient-bundling stress seller instead shows Both raising welfare by 18.7 and first-best attainment by 11.1 percentage points, while still lowering seller profit by 25.3 and raising buyer surplus by 44.0. The identical buyer-surplus gain masks opposite welfare signs.

  • The contract abstains selectively on the authors' own study. C2 is Invalid for the original estimate; C1 and C3 are Inconclusive for the controlled study; C4 passes. Neither "guardrails work" nor "guardrails fail" is licensed.

  • The paper does not claim guardrails are ineffective. It concludes that their apparent value is unidentified until the simulated agents and protocol pass the checks.

Methodology in Plain English

The authors use a configurable hotel transaction testbed. Each buyer profile has hidden component values (cancellation, quiet room, breakfast, view, high floor, late checkout), hard constraints, a willingness-to-pay cap, urgency, and an outside option. A seller offers a mandatory base set at a price plus separately priced optional additions. The buyer accepts the utility-maximizing feasible offer if it weakly beats the outside option.

All outcome metrics use assigned ground truth rather than values the agents state. Buyer surplus is value minus price minus the outside option; seller profit is price minus cost; welfare is the sum. All three are zero when no transaction occurs, which is the key identity: price discrimination alone moves surplus between buyer and seller and changes welfare only if it changes completion or composition. Welfare is reported as a fraction of a profile-specific first best.

The platform has two binary rules — an information rule blocking questions and responses about budget, WTP, urgency, and outside options, and a conduct rule restricting the mandatory base to essential components and making other components declineable add-ons. Their 2×2 crossing yields None, Info, Conduct, and Both.

Four experiments drive the audit. E1 re-runs the model ladder with a single shared schema and chooser in all cells. E2 selects the four best-performing 14B profiles post hoc and generates each profile-condition three times, then bootstraps profiles and runs an exact sign-flip placebo over all 16 profile-level sign assignments. E3 compares three seller prompts across three profiles with two generations each. E4 replaces the LLM seller with two deterministic scripted policies — a profit maximizer and an inefficient-bundling stress policy — across all 60 profiles.

Models are 4-bit Qwen2.5-Instruct at 1.5B, 3B, and 14B parameters, with the same model playing buyer and seller, two seller-question rounds, a structured offer, temperature 0.2, and the same 30 held-out synthetic profiles in every cell. The original run contains 180 LLM dialogues in None/Both; the controlled 2×2 contains 360. E1 uses 20,000-draw percentile bootstrap intervals resampling the 30 profiles.

Why This Matters

Impact on research. The paper turns a routine limitation into a design contract. It shows that a simulated role ("self-interested seller") is a manipulation that must be validated, not prose that can be assumed, and that a thin guardrail bundled with template, parser, or decision-logic changes is not a thin guardrail. Its identifiability argument is telling: if two traces produce the same observable but differ in the hidden property of interest, no downstream function of that observable can recover it — adding more aggregate metrics cannot repair missing instrumentation.

Real-world applications:

  • Marketplace platform policy. Before a platform ships a privacy or disclosure rule based on simulation evidence, the simulation must show that its sellers actually pursue profit and that the measured effect is not a schema artifact.
  • Agent benchmark and leaderboard design. Benchmarks that report a single aggregate score per agent inherit unstated assumptions about the constructs they measure; the contract's four gates offer a way to flag when an aggregate is not interpretable.
  • Repeated-run budgeting. The finding that generation residuals account for 49.9% of variation at temperature 0.2 quantifies why single-sample agent evaluations omit the error bars that matter.
  • Procurement and vendor claims. Buyers of agent evaluation services can ask for the nine-item checklist as a precondition for accepting a welfare or performance claim.

Industry relevance. Agent-to-agent commerce is an active deployment target, with buyer agents searching and negotiating and seller agents configuring offers and prices. The paper's central practical claim is that a marketplace guardrail may improve buyer surplus while reducing seller profit one-for-one without creating any welfare at all, and that joint reporting is required to tell those cases apart. Making Invalid and Inconclusive acceptable outputs prevents unsupported policies from being justified by precise-looking simulation numbers.

Future Directions

  1. Cross-family and frontier-model replication. All LLM results here use one instruction-tuned model family, three sizes, 4-bit quantization, and self-play, where the same model plays both roles. Cross-family pairing with role prompts held fixed is described as the cheapest direct test; it is not yet done.

  2. A properly powered confirmatory run. The repeated-generation analysis covers four profiles selected after observing extreme effects, so it diagnoses instability and winner's curse but cannot estimate a population treatment effect. A confirmatory run would require pre-specified profiles and at least three generations per profile-condition.

  3. Repairing the incentive manipulation. The seller-prompt manipulation did not move profit monotonically with only three profiles and two generations per prompt. A larger manipulation check is needed to determine whether prompt-defined roles can be made controllable, or whether roles should only be described behaviorally.

  4. Richer environments and populations. The synthetic hotel panel has 43 negative outside-option utilities among 60 profiles under the stated outside-option formula, so most participation constraints weakly bind and a null Info effect would be confounded by the opportunity structure. An optimized fixed menu also has a training-distribution advantage over zero-shot agents. The authors explicitly frame the contract as a minimum gate, not a complete theory of external validity, platform equilibrium, or human welfare.

Target Audience

Researchers who build or consume LLM agent market simulations, benchmark designers, and evaluation engineers responsible for agentic commerce testbeds. Also useful for applied causal-inference practitioners interested in construct validity outside the social sciences, and for product or policy teams who need to know which simulated marketplace results can and cannot support a deployment decision. Readers looking for evidence about whether marketplace guardrails work will not find an answer here: the paper deliberately abstains on the policy question and reports on the measurement problem instead.

Authors’ abstract

Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.

Read the original paper