Research
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
Overview Research area: Mechanistic interpretability for multilingual language models, grounded in causal abstraction and invariant causal prediction. Technical level: Advanced. The paper is written f

- arXiv
- 2512.24842
- Published
- 2025-12-31
- Authors
- Yanan Long
AI summary
Overview
Research area: Mechanistic interpretability for multilingual language models, grounded in causal abstraction and invariant causal prediction.
Technical level: Advanced. The paper is written for readers comfortable with structural causal models, interchange interventions, activation patching, and Bayesian uncertainty quantification.
Scope: The paper proposes triangulation, a falsifiable acceptance rule requiring necessity, sufficiency, and cross-environment invariance before a proposed internal circuit is accepted as a mechanistic explanation, and lays out (but does not yet run) a comparative protocol across four model families, three translation datasets, and four baseline methods.
What This Paper Is About
Multilingual language models perform well on average but behave unpredictably across languages, scripts, and cultures, and most analyses of their internal states are merely associational: they show where information is encoded, not whether that information causes behavior. The paper argues that a mechanistic explanation should only be accepted if it survives causal interventions and continues to hold across environments that perturb surface form while preserving meaning. To enforce this, the author defines triangulation as an acceptance rule applied on top of automatically discovered candidate circuits, rather than as a discovery method itself.
Key Contributions
-
A formal acceptance rule grounded in causal abstraction. Triangulation is cast as an approximate transformation score (Eq. 1) computed over a preregistered distribution of interchange interventions, with explicit necessity, sufficiency, and invariance criteria plus cue-only negative controls.
-
A protocol for constructing and auditing predicate-preserving reference families. This includes predicate checkers, reference-family quality scores aggregating checker agreement, human inter-rater reliability targeting κ ≥ 0.8, and stress tests that intentionally inject predicate violations (entity swaps, meaning flips) to confirm triangulation rejects when the predicate separation assumption fails.
-
A comparative experimental protocol spanning multiple model families (Gemma 3 in 1B and 4B variants; Llama 3.2 in 1B and 3B variants; Mistral 3 in its 3B variant; and MADLAD-400-3B-MT, a T5-based multilingual MT model covering over 450 languages), language pairs, task types, and four baselines, with explicit compute accounting.
-
An extension to multimodal settings, applying the same methodology to vision-language models by holding the image fixed and varying the prompt language/phrasing, and using the framework to separate visual, linguistic/cue, and spurious mechanism classes.
Main Findings
-
No empirical results are reported. The paper presents a protocol and formal framework; it does not report circuit acceptance rates, triangulation scores, benchmark numbers, or any experimental outcomes. The comparative protocol is proposed, not executed.
-
Triangulation reduces to interchange-intervention accuracy in a special case. When the similarity function is an indicator of agreement on a binary outcome, the triangulation score T_tri (Eq. 1) reduces to interchange-intervention accuracy as defined by Geiger et al.
-
Acceptance requires three properties simultaneously. Necessity (knockout ablation of the circuit must drop the task score by at least τ_N), sufficiency (predicate-swap patching must transfer the behavior in the predicted direction by at least τ_S), and invariance (predicate-matched patching across environments must change the score by at most ε, and cue-only patching must also stay within ε).
-
The object under test is a mechanism class, not a single circuit. Because recent multilingual circuit analyses show the same abstract computation can be implemented by a mixture of shared and language-specific components, the paper defines a mechanism class as an environment-indexed family {C_e} with optional translation maps, and requires min over environments of T_tri(e) ≥ η. This permits language-specific "adapter" subcircuits as long as predicate-swap effects transfer correctly and cue-only falsifiers do not transfer behavior.
-
Translated patching addresses off-manifold cross-lingual activation replacement. When base and source environments differ, naive value replacement may be off-manifold or type-incompatible due to tokenization or position mismatch, so a learned map T_{e_b←e_s} aligns source activations into the base circuit's coordinate system before replacement, subject to an on-manifold bound δ on allowed activation distortion. Knockout ablations (mean ablation, resampling, or zeroing) are treated as a special case where source activations come from a null distribution.
-
Uncertainty is quantified with a Beta–Binomial model. Given n sampled interventions with k successes, the posterior is Beta(a₀+k, b₀+n−k), with a₀=b₀=1 or the Jeffreys prior a₀=b₀=1/2 suggested; the threshold η is calibrated against placebo circuits (random circuits or circuits optimized for nuisance prediction) so that the placebo acceptance rate approximates a target α.
-
Four baselines are proposed for comparison: single-environment patching, causal mediation analysis, causal scrubbing, and ablation-based circuit testing using faithfulness scores from ACDC-style methods. The comparison centers on acceptance rates as a proxy for the false-positive versus over-conservatism trade-off.
-
Four limitations are acknowledged: triangulation identifies sufficient but not unique mechanisms because redundant circuits may also pass; reference-family quality is decisive and poor families can either admit spurious circuits or reject genuine ones; causal interventions at scale are computationally expensive; and LLM judges used to verify predicate preservation risk circularity if correlated with the model under study.
Methodology in Plain English
The paper treats multilinguality as controlled environment variation. The researcher fixes a predicate function π that extracts the semantic property relevant to the behavior under study, then builds a reference family of variants that all share that same predicate but differ in language, script, or style. Whether a candidate circuit genuinely mediates the behavior is then tested by intervening on it across those variants rather than in a single setting.
The proposed pipeline has three stages. First, discovery: for each environment in a reference family, run an automatic circuit discovery method (for example EAP-IG or position-aware circuit discovery) to propose an environment-specific circuit for the chosen task score, and in parallel learn candidate cue circuits (language or script predictors) to serve as explicit negative controls. Second, translation maps: using predicate-matched pairs within reference families, fit maps that align activations for translated patching, with strict hold-out to prevent leakage into acceptance evaluation. Third, acceptance: evaluate the resulting mechanism class under an intervention distribution that includes knockout, predicate-swap patching, stability tests, and cue-only falsifiers, accepting only if the score is high overall and across every base environment. Separating discovery from acceptance is intended to ensure no manual head-picking influences the results, and the same discovered circuits are evaluated under all methods for fair comparison.
The task side uses two settings. In translation with a fixed target language (for example French), the behavioral score is a logit margin for inclusive versus binary gender realization at the target-side gender-bearing locus (GBL) — the earliest decoding position where the target language commits to gender marking. In intra-lingual rewriting using mGeNTE, the score is a neutrality classifier or logit margin measuring whether the rewrite avoids unnecessary gender marking while preserving the predicate. Data sources are FairTranslate for English-to-French (2,418 sentence pairs, each underlying proposition appearing in multiple gender-marked variants, with the gender variants treated as different predicate values for swap interventions rather than as reference-family members), GLITTER for English-to-German (a multi-reference benchmark with post-edited translations supporting multiple gender-fair strategies, with strategy choice treated as the environment axis), and mGeNTE, built on Europarl, spanning English to Italian, Spanish, German, and Greek.
Why This Matters
The paper raises the evidential bar for mechanistic claims in multilingual models, where large aggregate gains can mask instability across languages, writing systems, and cultures and where model rankings can invert on language- or culture-specific subsets. By requiring that causal effects stay directionally stable and of sufficient magnitude across predicate-preserving environments, triangulation is designed to filter spurious circuits that pass single-environment tests. Framing the rule as a proxy task with empirical feedback aligns it with the pragmatic interpretability agenda, and the author connects it to a "science of misalignment" theory of change: if a model misbehaves, the framework helps determine whether the responsible circuit is language-specific (and therefore spurious) or language-agnostic (and therefore a genuine mechanism).
Real-world applications:
- Gender-fair machine translation: auditing whether a model's gender handling reflects the actual referent constraint rather than surface cues such as source language or script.
- Cross-lingual question answering: verifying that the same question asked in different languages is answered through the same underlying mechanism.
- Entity preservation and morphosyntactic agreement: checking that names, quantities, number, and tense marking transfer consistently regardless of surrounding linguistic context.
- Vision-language model auditing: holding an image fixed and varying the prompt language to isolate circuits that extract visual information from those that merely correlate with linguistic cues.
Industry relevance: The framework targets teams that build or deploy multilingual and multimodal models and need an objective, preregistered criterion for deciding whether an internal explanation is trustworthy. It bears on safety evaluation, model auditing, and translation system quality assurance, and it offers a concrete way to compare acceptance criteria across existing interpretability methods rather than relying on single-environment patch scores.
Future Directions
- Execute the proposed protocol. No experimental results exist yet; the model families, datasets, baselines, and intervention distributions described would need to be run to determine whether triangulation actually achieves a better acceptance-rate trade-off than the four baselines.
- Improve and stress-test reference families. The paper notes quality is decisive, especially for less-resourced languages, and recommends reporting measured degradation in quality beyond high-resource pairs and analyzing how robust triangulation remains as a function of that quality.
- Resolve the uniqueness and redundancy problem. Multiple circuits may satisfy triangulation when their effects are redundant, so triangulation identifies sufficient mechanisms but does not guarantee uniqueness.
- Reduce computational cost and circularity risk. Causal interventions at scale are expensive and prioritization strategies are advisable; separately, LLM judges used for predicate preservation can inflate apparent reference-family quality if they correlate with the model under study, so rule-based or human checks are preferred where feasible.
- Extend beyond translation. The paper points to cross-lingual question answering, morphosyntactic agreement, and entity preservation as settings that share the predicate-preserving environment structure.
Target Audience
Researchers and graduate students in mechanistic interpretability, multilingual NLP, and causal inference who are interested in formal standards for validating circuit-level claims; safety and evaluation teams at organizations deploying multilingual or vision-language models; and methodologists who want a falsifiable criterion they can compare against existing approaches such as single-environment patching, causal mediation analysis, causal scrubbing, and ACDC-style faithfulness testing. The paper assumes familiarity with structural causal models, activation patching, and causal abstraction, making it most suitable for an advanced audience.
Authors’ abstract
Multilingual language models achieve strong aggregate performance yet often behave unpredictably across languages, scripts, and cultures. We argue that mechanistic explanations for such models should satisfy a \emph{causal} standard: claims must survive causal interventions and must \emph{cross-reference} across environments that perturb surface form while preserving meaning. We formalize \emph{reference families} as predicate-preserving variants and introduce \emph{triangulation}, an acceptance rule requiring necessity (ablating the circuit degrades the target behavior), sufficiency (patching activations transfers the behavior), and invariance (both effects remain directionally stable and of sufficient magnitude across the reference family). To supply candidate subgraphs, we adopt automatic circuit discovery and \emph{accept or reject} those candidates by triangulation. We ground triangulation in causal abstraction by casting it as an approximate transformation score over a distribution of interchange interventions, connect it to the pragmatic interpretability agenda, and present a comparative experimental protocol across multiple model families, language pairs, and tasks. Triangulation provides a falsifiable standard for mechanistic claims that filters spurious circuits passing single-environment tests but failing cross-lingual invariance.