Skip to content
AI.info

Research

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

Overview Research area: Bias auditing and interpretability for large language models, specifically representation-level analysis of fine-tuning side effects. Technical level: Intermediate. The paper a

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
arXiv
2609.10060
Published
2026-09-09
Authors
Marek Jeliński, Jan Dubiński, Maciej Chrabaszcz, Sebastian Cygert

AI summary

Overview

Research area: Bias auditing and interpretability for large language models, specifically representation-level analysis of fine-tuning side effects.

Technical level: Intermediate. The paper assumes familiarity with hidden-state embeddings, cosine similarity, fine-tuning (full and LoRA), model merging, and bias benchmarks, but the core idea is explained with accessible geometric intuition.

Scope in one sentence: The paper introduces a reference-based auditing metric, the Representational Bias Shift (ΔB), which compares the hidden-state geometry of a fine-tuned model against its base model in a shared "relative representation" space, and tests whether that internal shift co-varies with output-level bias measured by three external benchmarks.

What This Paper Is About

Most bias auditing for LLMs looks only at what a model generates, which requires expensive curated benchmarks or judge models and can miss internal changes that never surface in text. The authors ask a narrower empirical question: when fine-tuning shifts a model's hidden-state associations with social groups, does that shift track the change in output-level bias? To answer it, they project both an audited (fine-tuned) model and a reference (base) model into a shared comparison space using relative representations, and define the difference in bias between them as ΔB.

Key Contributions

  1. A reference-based auditing framework that places an audited model and a reference model into a shared comparison space through relative hidden-state representations, defining the Representational Bias Shift ΔB as the change in how target groups associate with positive versus negative attribute sentences.
  2. Validation of ΔB against three output-level benchmarks (WildGuardMix, DecodingTrust, ToxiGen) across three model families (Mistral, Llama, Gemma) and two fine-tuning regimes (full and LoRA), using a graded merge spectrum so bias is introduced in increments rather than as a single jump. ΔB co-varies with output-level bias in 15 of the 18 settings tested, with |r| up to 0.84.
  3. Experiments validating relative representations against three alternatives (SEAT, Procrustes-SEAT, CKA drift), showing RR is the strongest method on the benchmarks tested, with Procrustes-SEAT staying near chance.
  4. Ablations showing ΔB is stable across anchor sets, attribute sets, target templates, pooling strategies, and independent fine-tuning runs.

Main Findings

  • Correlation with output-level bias: ΔB correlates with output-level bias change in 15 of the 18 settings tested. Under full fine-tuning the correlation reaches |r| = 0.84 (p < 0.001). The association is weakest for Gemma across all benchmarks and both regimes.

  • Detection of increased-bias checkpoints: Thresholding ΔB detects checkpoints whose bias increased with ROC AUC between 0.65 and 0.99. On WildGuardMix under full fine-tuning the ROC AUC is 0.93 for Mistral, 0.89 for Llama, and 0.78 for Gemma.

  • WildGuardMix (harmfulness): Full fine-tuning Pearson correlations are −0.67 for Mistral, −0.68 for Llama, and −0.37 for Gemma. Under LoRA they are −0.61, −0.65, and −0.04 respectively, with ROC AUCs of 0.78, 0.92, and 0.77.

  • DecodingTrust (stereotype agreement): Mistral and Llama correlate strongly under full fine-tuning (r = −0.82 and −0.84, p < 0.001), with detection strongest for Llama at ROC AUC 0.91. Gemma is weaker but still significant under full fine-tuning; its LoRA correlation falls below significance.

  • ToxiGen (per-group toxicity): Llama shows r = −0.62 under full fine-tuning, Gemma r = −0.31, and pooling all three families gives r = −0.49 (p < 0.001), with detection reaching ROC AUC 0.91 for Llama. Mistral is the exception under full fine-tuning (r = −0.19, p = 0.14), attributed to its generated toxicity saturating on the more harmful merged checkpoints. Under LoRA all three families are significant, including Mistral (r = −0.43, p < 0.001), with detection between 0.69 and 0.86.

  • Relative representations beat baselines: On Llama, RR stays above every baseline across the full threshold sweep on both benchmarks. SEAT reaches ROC AUC 0.778 against RR's 0.964 on WildGuardMix. Procrustes-SEAT sits at chance (0.5) on both benchmarks, so the gain comes from the relative representation rather than from alignment alone. CKA drift detects well (0.864 and 0.753) but stays below RR, so ΔB is not reducible to representational displacement.

  • Anchor choice: Neutral in-domain sentences perform best, reaching ROC AUC 0.892 at 1k anchors, ahead of the original word anchors and samples from Alpaca and the Tulu mixture.

  • Robustness to wording: Varying six attribute constructions and six target templates gives mean ROC AUC 0.863 ± 0.030 across attribute sets and 0.902 ± 0.008 across templates. Pooling strategy (mean, max, last) leaves the ordering unchanged, with mean pooling strongest. Repeated training with different random seeds yields nearly identical ΔB values.

  • Computational cost: The method needs roughly 3 minutes per model (2m51s for Llama, 3m12s for Mistral, 3m12s for Gemma, split between generating embeddings and computing the bias shift). Output-level benchmarks take 9m03s to 14m02s for WildGuardMix, 33m10s to 75m36s for ToxiGen, and 44m22s to 155m36s for DecodingTrust — between three and roughly fifty times more compute.

  • Gemma weakness: Gemma is the weakest case throughout. The authors note that Gemma-3-4B is roughly half the size of the Mistral and Llama models audited, that its correlations keep the same sign everywhere (so the signal is weak rather than reversed), and that they controlled for neither tokenisation nor final-layer geometry.

Methodology in Plain English

The core problem is that fine-tuning changes the geometry of a model's internal space, so you cannot directly compare the raw hidden states of a fine-tuned model against those of its base model — even the same sentence may land in a different region. The authors solve this by encoding each sentence not by its own embedding but by its cosine similarities to a fixed set of anchor sentences. Both models encode the same anchor sentences with their own parameters, so the resulting "relative representation" vectors live in a shared coordinate system without fitting any cross-model map.

Within that shared space, the authors take sets of target sentences (social-group related), positive attribute sentences, and negative attribute sentences. Because relative representation coordinates are already cosine similarities, they measure associations using negated Euclidean distance rather than cosine similarity again, and produce a mean bias score B_rel per model. The Representational Bias Shift is simply ΔB = B_aud − B_ref: a negative value means the target group moved toward the negative attributes, read as increased bias.

To test whether this proxy tracks real behaviour, they fine-tune Llama 3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Gemma 3-4B-IT on an unharmful and a synthetically harmful split of WildGuardMix, each of 8k examples, under both full and LoRA fine-tuning. They then linearly merge the two checkpoints at five interpolation ratios (10/90, 30/70, 50/50, 70/30, 90/10 of unharmful to synth), giving seven checkpoints per model and regime spanning safe to harmful. For each checkpoint and target group they pair ΔB with an external ΔBias Score from WildGuardMix (scored with the allenai/wildguard guard model), DecodingTrust's stereotype pipeline, or ToxiGen (scored with the authors' toxigen_roberta classifier), then pool the pairs and report Pearson correlations, ROC AUC of a threshold classifier on ΔB, and the mean absolute error of a linear fit estimated over 1,000 bootstrap resamples.

Why This Matters

Impact on research: The paper tests a contested assumption — that representation-level metrics predict downstream behaviour — in a specific, narrow setting (measuring fine-tuning-induced shifts rather than debiasing or single-model bias). It provides one of the more affirmative empirical answers, with the caveat that the association is weaker under parameter-efficient adaptation and for Gemma. It also gives a lightweight, anchor-based alternative to fitting explicit cross-model maps or comparing outputs on curated data.

Real-world applications:

  • Auditing fine-tuned model checkpoints before deployment, without building a benchmark for every new harm or target group.
  • Monitoring fine-tuning runs for unintended safety degradation, potentially as an early-stopping signal.
  • Comparing independently trained or versioned checkpoints, since the audited model does not need to originate from the reference model.
  • Rapid triage across many candidate checkpoints, given the roughly three-minute per-model cost.

Industry relevance: The method requires no task-specific evaluation data and uses 3 to 50 times less compute than the output-level benchmarks considered, making it practical for teams that fine-tune or adapt models frequently and cannot afford full benchmark suites at every checkpoint. The authors explicitly frame it as complementary to output-based auditing rather than a replacement, and note it is a detection and auditing tool rather than a mitigation method. They also flag dual-use risk: a cheap differentiable signal could be optimised against, and a model tuned to keep ΔB small need not be less biased in its outputs.

Future Directions

  • Using ΔB as a monitoring signal during fine-tuning, for example as an early-stopping criterion, while noting that turning it into a training objective is less straightforward.
  • Investigating why the signal degrades for Gemma, including the possible roles of model scale (Gemma-3-4B is roughly half the size of the other two), tokenisation, and final-layer geometry — none of which the authors controlled for.
  • Extending beyond English templates, the coarse single-axis groups of DecodingTrust, and the harms those sentence sets name, toward other languages and intersectional groups.
  • Testing whether the correlation with output-level bias holds at larger scale or under naturally occurring fine-tuning, and whether the relationship is causal. The authors also call for systematic study of pooling and layer selection.

Target Audience

Researchers and engineers working on LLM safety, bias evaluation, and interpretability; practitioners who frequently fine-tune or adapt models and need cheap pre-deployment auditing signals; and methodologists interested in whether representational metrics track behavioural change. Readers without background in hidden-state analysis or fine-tuning will find the geometry of the method harder to follow, though the authors keep the motivation and framing accessible.

Authors’ abstract

Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before and after fine-tuning. Because fine-tuning reshapes representation geometry, absolute hidden states are not directly comparable, so we encode each sentence by its similarities to a fixed set of anchor sentences, yielding relative representations in a shared comparison space. There we measure how target groups shift in their association with positive and negative attributes, a quantity we call the Representational Bias Shift $ΔB$. Across three model families and the WildGuardMix, DecodingTrust and ToxiGen benchmarks, $ΔB$ correlates with output-level bias change in 15 of the 18 settings we test, reaching $|r| = 0.84$ ($p &lt; 0.001$) under full fine-tuning and becoming more model-dependent under parameter-efficient adaptation. Thresholding $ΔB$ detects checkpoints whose bias increased with ROC AUC between $0.65$ and $0.99$, and on WildGuardMix and DecodingTrust it separates them better than a SEAT-based baseline for all three families. $ΔB$ is also stable under changes to the anchor set, attribute sets and target templates. Our method requires no task-specific evaluation data and audits a model in about three minutes, using $3$-$50\times$ less compute than the output-level benchmarks considered here. We view it as complementary to output-based auditing rather than a replacement for it.

Read the original paper