Research
REMISVFU: Vertical Federated Unlearning via Representation Misdirection for Intermediate Output Feature
Overview Research area: Machine unlearning in Vertical Federated Learning (VFL), specifically client-level federated unlearning combined with privacy-preserving distributed training. Technical level:

- arXiv
- 2512.10348
- Published
- 2025-12-11
- Authors
- Wenhan Wu, Zhili He, Huanghuang Liang, Yili Gong, Jiawei Jiang, Chuang Hu, Dazhao Cheng
AI summary
Overview
- Research area: Machine unlearning in Vertical Federated Learning (VFL), specifically client-level federated unlearning combined with privacy-preserving distributed training.
- Technical level: Advanced. The paper assumes familiarity with federated learning paradigms (HFL vs. VFL), split learning / SplitNN architectures, backdoor attacks, membership inference attacks, and gradient-projection techniques from multi-task learning.
- Scope: The paper proposes ReMisVFU, a plug-and-play representation-misdirection framework that erases a departing party's contribution from a trained splitVFL model in a few epochs while preserving the utility of the remaining parties.
What This Paper Is About
Federated unlearning lets a participant withdraw from a trained federated model, as data-protection laws such as the GDPR and the California CCPA require. Nearly all existing unlearning methods target Horizontal Federated Learning, where different parties hold different samples with the same features, and they break down in Vertical Federated Learning, where parties hold different feature columns for the same users and the model itself is split across parties. The paper's goal is a fast, client-level unlearning method for splitVFL that removes the forgetting party's statistical influence, keeps the remaining parties' accuracy high, and requires no changes to the existing communication protocol.
Key Contributions
- Representation misdirection mechanism. The forgetting party collapses its bottom-model output features onto a randomly drawn anchor vector on the unit sphere, severing the statistical link between its features and the global model.
- Gradient coordination strategy. The server orthogonalizes the retention gradient against the forgetting gradient via projection, so the two objectives do not destructively interfere during joint optimization.
- Plug-and-play unlearning pipeline. The method reuses the exact communication pattern of standard VFL and leaves the workflow of the retained passive parties unaffected, so it can be integrated into existing splitVFL systems without protocol modification.
- Comprehensive evaluation. Experiments on five image benchmarks against four baselines show strong forgetting and retention performance, plus ablations and parameter-sensitivity analysis.
Main Findings
- Backdoor suppression: ReMisVFU drives backdoor attack success rate close to the natural class prior. Reported backdoor rates on the five benchmarks are 9.96 (MNIST), 10.39 (Fashion-MNIST), 10.22 (SVHN), 10.55 (CIFAR-10), and 18.33 (CIFAR-100), compared with the "Original (no Bkd.)" priors of 9.06, 9.91, 9.86, 9.63, and 1.91 respectively.
- Clean-accuracy cost: ReMisVFU's clean accuracy is only 2.45% lower than the fully retrained gold model on average (the abstract states approximately 2.5 percentage points). By comparison, FedR2S, FedGA, and FedKD are on average 5.10%, 3.83%, and 4.22% lower than the retrained model.
- Comparison to FedR2S: FedR2S leaves substantial backdoor residue, increasing it by an average of 7.84%, indicating that switching optimizers after loading a pretrained model is insufficient for thorough forgetting.
- Comparison to FedGA and FedKD: FedGA suppresses triggers to roughly 9–14% on easier datasets but deteriorates to 22.1% on CIFAR-100; FedKD's performance falls between FedR2S and FedGA.
- Membership inference privacy: ReMisVFU attains MIA metrics closest to FedRetrain on every dataset. The average AUC gap to FedRetrain is 4.22% (5.82% for ACC). FedKD and FedR2S show average AUC gaps of 12.28% and 6.50% (ACC gaps of 8.89% and 5.94%). FedGA scores 1.000 AUC and 1.000 ACC on all five datasets, which the authors attribute to its gradient-ascent step driving the model into a highly over-fitted, sharp region of the loss landscape.
- Runtime efficiency: ReMisVFU completes forgetting using only 10.63%–21.49% of FedRetrain's time, and cuts wall-clock time by 24.33% relative to FedR2S. FedGA requires roughly 3.89 times ReMisVFU's runtime, and FedKD roughly 1.77 times.
- Ablation on gradient coordination: Removing the gradient conflict-mitigation strategy (No-GCM, vanilla gradient summation) raises the backdoor attack success rate by 2.47% and lowers clean accuracy by 1.71% on average. Replacing orthogonal projection with a random unit vector of the same norm (Rand-Proj) causes pronounced degradation in both forgetting effectiveness and retained utility — for example 35.62 clean accuracy and 82.63 backdoor rate on SVHN.
- Parameter sensitivity: For the anchor scaling factor c, both clean accuracy and backdoor success remain essentially flat when c ≤ 2, but forgetting degrades sharply once c ≥ 4 because a large anchor pushes features into the saturation regime of activations. For the trade-off coefficient α, clean accuracy rises slightly and backdoor success stays near random guessing for 1×10⁻⁴ ≤ α ≤ 1×10⁻³, while backdoor success exceeds 50% once α surpasses 5×10⁻³.
- Representation-level evidence: t-SNE visualizations on MNIST and Fashion-MNIST show that after only a few iterations the unlearned feature distribution closely matches the fully retrained gold baseline, but with more uniform density.
Methodology in Plain English
In vertical federated learning, several organizations each hold different features about the same people. Each organization runs a small local encoder ("bottom model") on its own features, sends the intermediate output features to a label-holding active party, which runs a top model that produces the prediction. When one organization asks to be forgotten, you cannot simply delete rows of data, because its influence is spread across both its own encoder and the shared top model.
ReMisVFU handles this at the representation level. Before unlearning begins, the system samples one random unit vector on the unit sphere and fixes it for the whole procedure. During unlearning, the forgetting party's encoder is pushed so that its output features for every input collapse onto that same fixed anchor vector, scaled by a hyperparameter c. If every input produces essentially the same output vector, the representation carries no usable information about that party's data, so the top model can no longer recover anything from it.
At the same time, the model must keep working for the remaining parties, so the server also computes the ordinary task loss (cross-entropy for classification, MSE for regression) on the retained data. The problem is that the gradient that improves retention and the gradient that forces forgetting often point in opposite directions; naively adding them lets them cancel out, which makes forgetting slow and unstable. The paper's fix is a projection step: compute the cosine similarity between the two gradients, and if it is negative, subtract from the retention gradient its component along the forgetting gradient, leaving a residual that is orthogonal to the forgetting direction. The forgetting gradient is used in full; only the non-conflicting part of the retention gradient is added. The whole loop otherwise reuses the standard VFL forward and backward pass, so no protocol changes are needed.
The evaluation partitions each image into three equal vertical slices (left, centre, right), assigning the centre slice to the target party to be forgotten. Each client encoder has two 3×3 convolutional layers (32 and 64 channels) each followed by 2×2 max-pooling, and the top module concatenates all client features and applies two fully connected layers with ReLU hidden size 128. Backdoors are implanted with a 2×2 white trigger stamped in the lower-right corner with target label "0", poisoning about 10% of the target party's images.
Why This Matters
Impact on research. The paper addresses a gap it identifies explicitly: almost all federated unlearning research targets horizontal federated learning, and the existing vertical federated unlearning work is limited to logistic regression models, gradient-boosted decision trees, checkpoint-based rapid retraining (which incurs significant storage overhead), or knowledge distillation that changes architecture rather than truly removing memorized information. ReMisVFU instead intervenes on internal representations, positioning unlearning as a representation-editing problem rather than a retraining or output-matching problem.
Real-world applications:
- Cross-industry credit scoring, where an e-commerce platform holding purchase histories and a bank holding income and credit scores train a joint model. If one party withdraws, its feature contribution must be removed without the other party losing its model.
- Healthcare consortia where a hospital supplies labels or diagnoses while other institutions supply disjoint feature sets; withdrawing an institution must not invalidate the deployed model.
- Financial or telecommunications collaborations subject to GDPR or CCPA erasure requests, where retraining the joint model from scratch is prohibitively expensive.
- Any splitVFL deployment that requires fast, low-overhead compliance responses, since the method avoids protocol changes and reuses existing communication patterns.
Industry relevance. The runtime results matter for deployment: using 10.63%–21.49% of retraining time means an erasure request can be honored in a small fraction of the cost of rebuilding the model, and 24.33% faster than FedR2S. The plug-and-play property means existing splitVFL infrastructure does not need to be redesigned. The finding that backdoor success returns to roughly the natural class prior is directly relevant to supply-chain and insider-threat concerns in federated deployments.
Future Directions
- Anchor scaling beyond the tested range. The paper shows forgetting collapses once c ≥ 4, attributing this to activation saturation. A principled way to choose c without a per-dataset sweep, or an adaptive anchor magnitude, is an open question the sensitivity analysis raises.
- Extension beyond classification. The retention loss is defined for cross-entropy or MSE, so regression is nominally supported, but all reported evaluations are classification on image benchmarks. Behavior on tabular or non-image VFL data is not reported.
- Active-party erasure. The paper deliberately excludes erasure of label data because removing labels would strip the supervisory signal and preclude further classification. Handling forgetting requests from a label-holding party, or label-only settings where d_K = 0, remains unaddressed.
- Stronger privacy auditing. The MIA evaluation uses a binary XGBoost classifier over output logits. Whether representation misdirection resists stronger or adaptive attacks that specifically target the collapsed-anchor structure is not reported.
Target Audience
Researchers and practitioners working on federated learning, machine unlearning, and privacy-preserving machine learning — particularly those building or operating vertical federated (splitVFL) systems that must satisfy data-deletion regulations. The paper is also relevant to readers interested in representation-level model editing, gradient-conflict resolution in multi-objective optimization, and backdoor mitigation in distributed training.
Authors’ abstract
Data-protection regulations such as the GDPR grant every participant in a federated system a right to be forgotten. Federated unlearning has therefore emerged as a research frontier, aiming to remove a specific party's contribution from the learned model while preserving the utility of the remaining parties. However, most unlearning techniques focus on Horizontal Federated Learning (HFL), where data are partitioned by samples. In contrast, Vertical Federated Learning (VFL) allows organizations that possess complementary feature spaces to train a joint model without sharing raw data. The resulting feature-partitioned architecture renders HFL-oriented unlearning methods ineffective. In this paper, we propose REMISVFU, a plug-and-play representation misdirection framework that enables fast, client-level unlearning in splitVFL systems. When a deletion request arrives, the forgetting party collapses its encoder output to a randomly sampled anchor on the unit sphere, severing the statistical link between its features and the global model. To maintain utility for the remaining parties, the server jointly optimizes a retention loss and a forgetting loss, aligning their gradients via orthogonal projection to eliminate destructive interference. Evaluations on public benchmarks show that REMISVFU suppresses back-door attack success to the natural class-prior level and sacrifices only about 2.5% points of clean accuracy, outperforming state-of-the-art baselines.