Skip to content
AI.info

Research

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

Overview Research area: Machine learning safety and security — specifically tamper resistance (preventing fine-tuning attacks from removing safeguards) in open-weight foundation models, with a theoret

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models
arXiv
2610.09004
Published
2026-10-06
Authors
Domenic Rosati, Alessa Carbo, Ali Dadsetan, Hong Huang, Matthew Young, Subhabrata Majumdar, Frank Rudzicz, Hassan Sajjad

AI summary

Overview

Research area: Machine learning safety and security — specifically tamper resistance (preventing fine-tuning attacks from removing safeguards) in open-weight foundation models, with a theoretical treatment of mutual-information-based safety certification.

Technical level: Advanced. The paper is built on mutual information, group actions, Jacobians, tangent kernels, and Polyak–Łojasiewicz convergence analysis, together with controlled empirical studies.

Scope: The paper argues theoretically and demonstrates empirically that mutual information measured at model release cannot by itself certify that an adversary will need many fine-tuning steps to recover a removed capability.

What This Paper Is About

When a lab removes harmful data or unlearns a dangerous capability from a model and then releases its weights, a natural question is whether the removal guarantees that an attacker needs many updates to bring the behavior back. A common assumption is that if the harmful information is gone — measured by mutual information — recovery must be slow. This paper shows that this assumption fails: information content at release says nothing about the geometry of the attack, so models with identical (even zero) information can have wildly different recovery times after fine-tuning.

Key Contributions

  1. A general decoupling result. Lemma 3.1 shows that function-preserving symmetries (such as positive rescaling in homogeneous networks or query–key rescaling in Transformer attention) leave a model's mutual information exactly unchanged while transforming its parameter Jacobian. Proposition 3.2 shows these symmetries act as different preconditioners on gradient descent, and Corollary 3.3 shows any information-based certificate is capped by the fastest representative in the symmetry orbit.

  2. Application to training-data filtering. Corollary 4.1 shows that complete exclusion of an independently drawn harmful dataset yields zero weight–data mutual information without constraining the optimization geometry. Proposition 4.2 shows that, at fixed information, changing the position of a retained harmful sample in the training stream changes the recovery time.

  3. Application to capability removal (unlearning). Proposition 5.2 constructs a release with perfect label–representation independence, I(Y; Z) = 0, that preserves the entire pretrained parameter Jacobian and tangent kernel. Proposition 5.3 gives an explicit zero-information release that recovers in a single gradient step for every 0 < ε < 1/2.

  4. Empirical demonstrations. On the Deep Ignorance 6.9B strong-filter (SF) and unfiltered (UF) checkpoints of O'Brien et al. (2026) attacked on WMDP-bio, a query–key gauge rescaling removes and even reverses the recorded filtration advantage. A BERT-Tiny scrubbing experiment shows an exact Q/K transformation reducing recovery from 280 to 25 steps with the same released function.

Main Findings

  • Information removal does not lower-bound recovery time. An explicit construction has both weight–data mutual information and label–representation mutual information equal to zero and recovers in one gradient step.

  • Parameterization, not information, can drive apparent tamper resistance. After 293 SGD updates on WMDP-bio with three attack seeds, SF at α = 8 reaches 69.3 ± 4.4% training accuracy versus 68.9 ± 1.8% for UF at α = 1; unscaled SF reaches 58.5 ± 5.0%. Rescaling UF to α = 1/24 moves its median first attainment of 60% training accuracy from 125 to 225 updates, compared with 150 for unscaled SF and 125 for rescaled SF. The transform requires a single pass over the Q/K tensors with no additional training or data.

  • The advantage can reverse, not just vanish. Rescaling UF to α = 1/24 lowers its final held-out accuracy from 60.0 ± 2.2% to 42.7 ± 5.4%, below unscaled SF's 45.3 ± 4.1%. Against rescaled SF at α = 8, the comparison is 69.3 ± 4.4% training and 54.0 ± 5.0% held-out accuracy versus slowed UF's 61.5 ± 7.6% and 42.7 ± 5.4%.

  • Training order changes recovery at fixed information. In the scalar construction, I(𝒟_H; Θ_i) = H(Y) for every sample position i, yet the hitting time depends on position because the initial loss ℓ(Θ_i, Y) = ½(1 − c_i)² with c_i = η(1 − η)^{T−i−1} varies. Later safe updates silently attenuate a retained sample's behavioral effect while its label stays exactly recoverable from the sign of the parameter.

  • Order effects appear in language models. With Pythia-160M on AG News classification and E2E-NLG generation, recovery time trends downward as target data is introduced at a later position in a fixed-length utility-training stream. The paper notes that exact equality of mutual information across orders is established by the construction, not measured in the language models.

  • Zero information with an intact Jacobian is possible. Proposition 5.2 releases W* = 0, b* = u, which globally minimizes the removal objective and satisfies I(Y; Z) = 0, while its parameter Jacobian D_{(W,b)} Z_{W,b}(x) = [I ⊗ φ(x)ᵀ I] is unchanged, so its tangent kernel on fixed evaluation data is also unchanged.

  • The chosen minimizer matters. For Z_{a,b}(X) = abX, both (a, b) = (0, 1) and (0, 0) globally minimize ½𝔼[Z²] and have zero label mutual information. One unit gradient step from (0, 1) reaches (1, 1) and zero loss, while (0, 0) has zero Jacobian and remains stationary.

  • Tangent scale sets attack speed at identical information. In the factorized-head experiment on a frozen Pythia-160M backbone (z = A(s B_source h), released at A = 0), all scales have I(Y, Z) = 0, while J_A^(s) = s J_A^(1) and K_A^(s) = s² K_A^(1). For s > 0, the reparameterization à = sA makes SGD at rate η equivalent to the s = 1 attack at effective rate ηs².

  • Kernel alignment bounds recovery from above. Proposition 5.4 shows that PL-type conditions imply L(θ_t) − L* ≤ (1 − ημ)^t Δ_0 with μ = αλ, giving a sufficient-step count T_suff(δ) = ⌈log(Δ_0/δ) / (−log(1 − ημ))⌉. The authors stress this is an upper bound on attack time, not a lower-bound certificate of resistance.

  • Recovery accelerates after a real unlearning intervention. In the BERT-Tiny scrub experiment, target-adapted training gives 91.7% target accuracy, an estimated predictive 𝒱-information of 0.457, and recovery in 0 steps. MSE scrub + retain gives 53.4 ± 3.6% target accuracy, 0.115 estimated 𝒱-information, and 252 ± 67 recovery steps over three seeds. A paired seed-0 comparison shows an exact Q/K transformation reducing recovery from 280 steps at α = 1 to 25 steps at α = 32, with the same released function and representations.

  • Controls behave as expected. The no-target-adaptation control sits at approximately 50% target accuracy with 75 ± 5 recovery steps, and random initialization gives 50.0% target accuracy, 0.140 estimated 𝒱-information, and more than 1000 recovery steps.

Methodology in Plain English

The authors combine formal argument with controlled experiments.

Theory. They define recovery time as the first optimization step at which the attacker's expected loss drops below a threshold. They then identify transformations of the weights that leave the model's outputs exactly the same — such as multiplying one weight matrix by a constant α and another by 1/α — but change how large the gradients are in each coordinate. Since these transformations do not change the model's function, they cannot change any information measure computed from that function; but they change the effective learning rate, and therefore how fast an attack proceeds. Any certificate that depends only on information must be valid for the fastest model in this family of equivalent parameterizations, which caps how strong the certificate can be.

They then apply this to two settings. First, filtering harmful data: excluding a freshly drawn harmful dataset makes the released weights statistically independent of that draw, but says nothing about what the model already knows or how easily it learns the task. They show with a simple scalar model that just moving a retained sample earlier or later in the training stream changes recovery time while keeping the information identical. Second, unlearning: they build a representation that is provably independent of the target label — so no decoder can beat the label-only baseline — yet whose parameter derivatives are completely intact, and in one case recovers after a single gradient step.

Experiments. They take an existing data-filtering comparison (Deep Ignorance 6.9B strong-filter versus unfiltered checkpoints) and apply a query–key rescaling to each checkpoint that provably preserves its released function, then re-run the same fine-tuning attack. They run order-variation experiments on Pythia-160M with AG News classification and E2E-NLG generation, a factorized-head experiment on a frozen Pythia-160M backbone across tangent scales, and a post-unlearning experiment on BERT-Tiny with and without a Q/K transformation. Recovery times are reported with seed counts and standard deviations where multiple seeds were used.

Why This Matters

Impact on research. The paper challenges a widely used intuition in AI safety: that removing information is equivalent to making removal hard to undo. It shows that information content and optimization geometry are separate axes, and that evaluations of tamper resistance can be invalidated by a preprocessing step that costs one pass over the weights. It also reframes the certification problem — what is missing is not a better information measure but constraints on attack dynamics.

Real-world applications:

  • Safety evaluation of released open-weight models. Red teams and model auditors must now account for cheap, function-preserving reparameterizations before concluding that a filter or unlearning step provides durable protection.
  • Pretraining data governance. Decisions about filtering harmful pretraining data cannot be justified by an information-removal argument alone; the paper's Deep Ignorance analysis indicates the measured advantage can be reproduced or reversed in the unfiltered model.
  • Machine unlearning compliance. Organizations relying on unlearning to satisfy data-removal or capability-removal requirements need checks on the residual parameter geometry, not just decoder-estimated information.
  • Attack-cost forecasting. The tangent-kernel and PL analysis offers a route to predicting — as an upper bound — how many fine-tuning steps an adversary needs.

Industry relevance. Model providers that release weights, and the enterprises that deploy them, depend on claims about how hard safeguards are to strip. The paper shows that coordinate choices in a checkpoint can determine the outcome of an attack comparison, which affects benchmark design, model cards, and any downstream risk assessment built on those benchmarks.

Future Directions

  • Bound progress rather than lower-bound time. The authors state that a genuine resistance certificate must bound attainable progress throughout allowed attack trajectories, not just upper-bound attack time as their PL analysis does.

  • Account for reparameterization-invariant attacks. Open questions concern attacks such as natural gradient that are not sensitive to the coordinate choices exploited here, and how certification behaves under them.

  • Handle feature learning and kernel evolution. A kernel statistic measured only at release does not control behavior when features change during training; the evolution of K_{θ_t} needs to be characterized.

  • Find the right geometric or information-theoretic quantity. The paper raises whether Fisher information, which measures local parameter sensitivity rather than statistical dependence between random variables, could support a global notion of information content — and notes this remains open. It also points to global constraints such as barriers or mode connectivity, and notes that inference-time information bounds are left intact by these training-time results.

Target Audience

This paper is most useful to AI safety and alignment researchers working on tamper resistance, machine unlearning, and pretraining data filtering; to machine learning theorists interested in optimization geometry, symmetry groups, and information-theoretic generalization bounds; to red-team and evaluation practitioners designing fine-tuning attack benchmarks; and to policy and governance audiences who need to understand the limits of information-removal arguments when open-weight models are released. Readers will need comfort with mutual information, Jacobians, and gradient-descent convergence analysis to follow the formal sections, though the empirical findings on the Deep Ignorance and BERT-Tiny checkpoints are accessible with less mathematical background.

Authors’ abstract

Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.

Read the original paper