Research
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Overview Research area: Natural Language Processing, specifically multilingual evaluation of tool-using (agentic) large language models. Technical level: Advanced. The paper is built around a statisti
- arXiv
- 2608.11110
- Published
- 2026-08-11
- Authors
- Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
AI summary
Overview
Research area: Natural Language Processing, specifically multilingual evaluation of tool-using (agentic) large language models.
Technical level: Advanced. The paper is built around a statistical estimand, bootstrap intervals, permutation-based chance floors, variance decomposition, and pre-registered causal interventions; the argument is as much about measurement validity as about multilingual behaviour.
Scope in one sentence: The paper makes the executed action trace — not the final answer — the measured object of multilingual evaluation, across 8 models, 6 parallel benchmarks, 41 languages, 505 cells and 2,382,875 agent rollouts, and asks whether a tool-using agent takes the same route when the same task is posed in a different language.
What This Paper Is About
Multilingual evaluation of language models almost always compares final answers, discarding the intermediate actions an agent takes. For an agent, however, the route is the product: it determines cost and latency, it decides how the system fails, and it is the only part of behaviour that can be audited or governed through tool permissions and escalation rules. The goal here is to define and measure cross-lingual policy retention — the share of a model's own self-consistency that survives a change of language — and to show that the naive way of computing it is not merely noisy but unidentifiable, because five separate confounds sit between raw trace similarity and any defensible claim.
Key Contributions
- A ceiling-corrected estimand. Normalised policy retention, Ĩ = I_cross / I_within, pairs each model against its own same-language reproducibility rather than against an assumed ideal, making the quantity identifiable given that models are not self-consistent even within one language.
- A measurement protocol that prices five confounds on its own data, two of them established causally rather than assumed (trace length and the chance floor), and every one of which turns out to have been suppressing the effect rather than creating it.
- A scale-and-vendor extension that locates the boundary of the regularity — below roughly 10B parameters the frontier convergence breaks down, and the apparent ordering among smaller models is shown to be an artifact of the measured chance floor.
- A causal account of the mechanism (the English pivot), including a pre-registered prediction confirmed across four models, plus the demonstration that a single trace-extraction regex, not the model, manufactured an apparent multilingual failure.
Main Findings
- Cross-lingual policy divergence is real and structural. Under greedy decoding the gap is positive in all 24 cells, with every task-level bootstrap interval on the raw gap excluding zero. Across the temperature ladder T ∈ {0, 0.3, 0.5, 0.7, 1.0}, cross-lingual agreement is essentially flat while self-consistency falls sharply — the paper concludes divergence is invariant to sampling temperature.
- Every correction makes the effect larger. The uncorrected, unmatched baseline reports +0.0625; matching the token budget raises it to +0.0878; strict empty-trace exclusion with a fourth model gives +0.0861; and on the identical models and tasks at T = 0 the length-matched gap rises to +0.2074. The matched estimate is 38% larger than the unmatched one.
- Four frontier models converge on nearly the same retention. Under greedy decoding, Gemma-3-27B, Sarvam-M, Qwen3-235B and Llama-4-Maverick each keep 71–73% of their action policy when the language changes, landing within 2.6 percentage points of each other, with every interval narrower than the band itself.
- Model identity explains almost none of the variance at T = 0. Model identity accounts for 5.7% of the variance in Ĩ at T = 0 versus 74.8% at T = 0.5; the benchmark explains 26.9% versus 6.9%; residual is 67.4% versus 18.3%. The relative spread falls from 16.1% at T = 0.5 to 3.5% at T = 0.
- Uncorrected per-model rankings are a ranking of determinism. The correlation between a model's own self-consistency and its measured gap is r = +0.43 at T = 0.5 and r = +0.97 at T = 0. Ranking models by the T = 0.5 gap versus the T = 0 gap gives rank correlation ρ = −0.80 on identical models and tasks.
- The regularity has a boundary below roughly 10B parameters. At least two of the three added models (Gemma-3-4B, Qwen3-8B, Aya-Expanse-8B) fall outside the band on every way the authors computed it. The sharpest evidence is a within-family comparison: same recipe, 6.75 times the parameters, ten points apart. Correcting for the measured chance floor moves the 4B model inside the frontier band while the 8B falls below it — reversing which small model looks better.
- The chance floor is measured, not assumed (c ≈ 0.56). One 8B model, which emits the shortest traces in the study, scores 0.669 on two traces answering different tasks, and 0.947 on one benchmark where its cross-language agreement is 0.9497. Correcting for this floor lowers the retention level roughly fivefold while preserving the absolute band.
- Agents route non-English tasks through English. Translate is the most-used tool in every adapted benchmark for every compliant model, while the arithmetic-composition synthetic benchmark correctly inverts to Calc. Reasoning text is approximately 99% ASCII even on Devanagari, Tamil or Odia input. Reliance is uneven: Qwen3 spends 50.5% of its tool calls on Translate against Sarvam's 18.0%.
- The pivot is causally load-bearing. Removing Translate lowers length-matched cross-lingual agreement in proportion to how much a model uses it, and removal causes substitution rather than omission (Gemma into Search/Calc, Qwen3 and Llama-4 into Summarize, Sarvam into Search). Mandating Translate helps monotonically in head-room at 5/6, 5/6, 3/6 and 2/6 benchmarks, predicted in advance. Models will not abandon the English pivot when instructed to: asked to reason in the task's language, Gemma writes a non-Latin-script thought on 0.79% of non-Latin-script rollouts and Sarvam on 0.08%, a compliance rate under 1%.
- Trace length is a first-order driver, not a nuisance. A three-level manipulation holding task, language, seed and budget fixed moves mean trace length roughly threefold and moves Ĩ by 6–7 points with disjoint intervals — further than the entire band across the four frontier models. Across six protocol-adherent models, length explains under a third of the variance in retention.
- Measurement failure masqueraded as multilingual failure. GPT-OSS-120B yields no parseable trace on 76.4% of its rollouts and scores under 2% accuracy, yet posts the study's two highest raw invariance scores because empty pairs score 1.0. Two worked exemplars raise its measured accuracy twenty-sixfold while accuracy on readable outputs barely moves. Aya-Expanse-8B emits no parseable trace on 31.2% of rollouts; the authors recommend treating any model above approximately 20% as unranked.
- Accuracy and invariance are separate. Cross-lingual accuracy spread reaches 0.66 within a single model and benchmark (Qwen3: 0.990 on its best XCOPA language, 0.330 on its worst; Sarvam: 0.278 to 0.789 on XQuAD). Every model has an English advantage, and Sarvam-M, the Indic-specialised model, has the largest at +0.155. The pooled correlation of +0.897 between invariance and accuracy is an artifact of two clusters and falls to +0.378 once GPT-OSS is excluded.
- A correction that hurts is also reported. Self-consistency voting, measured with ten replicates, raises raw cross-language agreement but raises same-language agreement more, costing 1.6–1.9 points on the ceiling-corrected estimand.
Methodology in Plain English
Every task is rendered in 41 languages and pushed through one fixed symbolic tool-use scaffold with a shared five-tool alphabet; the model emits a Thought/Action trace plus an answer, and only the trace is compared. Tools are symbolic — calls are parsed and compared but never executed — so the study measures induced tool-use policy rather than grounded execution. The suite is six benchmarks on verified alignment keys: five repurposed public benchmarks (FLORES-200, XQuAD, XNLI, Belebele, XCOPA) plus a purpose-built synthetic benchmark, totalling 5,776 tasks, 41 languages across 17 families and 16 scripts, and 809 (benchmark, language-pair) comparisons over 428 distinct pairs. This alignment requirement is stricter than it looks: Bitext pairs each language with English but not with each other, so five screened corpora fail outright and a sixth stores its per-language configurations in different row orders.
The key design move is generating every cell twice under identical decoding, task set and token budget, varying only the serving seed. Same-language agreement (I_within) pairs the two replicates within one language; cross-language agreement (I_cross) pairs them across two. Because both sides are cross-seed, decoding noise enters identically and language is the only difference. Every comparison is length-matched in both directions and accepted only when both directions agree in sign; pairs are dropped whenever either trace is empty; intervals are task-level bootstraps. The central quantity is the ratio of the two. Default decoding is T = 0.5, max_tokens = 4096 and an iteration cap of 10, served with vLLM on 8× B200 nodes. Three checks precede every claim: language routing verified from prompt logs (prompt sharing between arms is 0.0%), truncation at 4096 tokens at most 0.63% for every model, and exact coverage across all 44 arms of the audited campaign. The main suite contains five instruction-tuned models — Gemma-3-27B, Sarvam-M (24B, Indic-specialised), Qwen3-235B-A22B, Llama-4-Maverick (17B active, 128 experts) and GPT-OSS-120B — with Gemma-3-4B, Qwen3-8B and Aya-Expanse-8B added to probe the boundary, giving eight models in total.
Why This Matters
Impact on research. The paper argues that multilingual evaluation of agents has been measuring the wrong object: answer-level parity is compatible with any amount of behavioural divergence. It supplies a corrected estimand and a protocol that prices its own confounds, and it shows that a widely used class of harness — single-pattern ReAct-style trace extractors — can manufacture apparent multilingual failures. It sits between two literatures that rarely meet: English-heavy agentic benchmarks that score outcomes, and cross-lingual suites that compare final predictions.
Real-world applications:
- Deploying multilingual agents where cost and latency budgets differ by language, since a translation-first route spends tokens the English arm never spends.
- Auditing and governance, where tool permissions, rate limits, escalation rules and audit trails are written against an expected action sequence — a policy validated on the English trace does not describe what the system does in Hindi.
- Regression testing and failure analysis, since an error introduced by a translation step exists only on the non-English path and will never surface in an English-only test.
- Evaluation tooling itself, where reporting parse-failure rates alongside headline numbers and treating models
Authors’ abstract
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.