Research
Private and interpretable clinical prediction with quantum-inspired tensor train models
Overview Research area: Privacy and interpretability of clinical machine learning, specifically membership inference attacks against clinical prediction models and a quantum-inspired tensor network de

- arXiv
- 2602.06110
- Published
- 2026-02-05
- Authors
- José Ramón Pareja Monturiol, Juliette Sinnott, Roger G. Melko, Mohammad Kohandel
AI summary
Overview
Research area: Privacy and interpretability of clinical machine learning, specifically membership inference attacks against clinical prediction models and a quantum-inspired tensor network defense.
Technical level: Intermediate. The paper combines applied machine learning privacy (membership inference, differential privacy) with tensor network methods from quantum many-body physics; readers need some familiarity with model access levels and low-rank decompositions, but the empirical setting is clinical and applied.
Scope: The paper attacks a deployed logistic regression model (LORIS) and shallow neural networks trained on immunotherapy response prediction, then proposes tensorizing discretized models into tensor trains (TTs) as a post-hoc defense that also adds interpretability.
What This Paper Is About
Publicly released clinical prediction models can leak whether specific patient cohorts were used in their training, and transparent models like logistic regression are especially exposed. The authors attack LORIS, a publicly deployed logistic regression on a U.S. government website, and show that cohort membership can be recovered with high confidence even from its public interface. They then propose a defense: converting trained models into tensor train representations built from discretized model outputs, which hides parameters, reduces black-box leakage, and preserves—and extends—interpretability.
Key Contributions
-
A real-world attack on a deployed clinical model. The authors attack LORIS, a publicly available logistic regression for immune checkpoint blockade (ICB) immunotherapy response prediction hosted on a U.S. government website, and recover both its parameters and its training cohort from queries through the public web interface.
-
A cohort-level membership inference framework with three access levels. They design shadow-model-based attacks under binary black-box (bBB), continuous black-box (cBB), and white-box (WB) access, applied to both logistic regression (LR) and shallow neural network (NN) target models, with a multi-label adversarial meta-classifier predicting which of six public cohorts were in the training set.
-
A quantum-inspired tensor-train defense. Using TT-RSS, they tensorize models from their discretized evaluations into tensor trains (with discretization into b bins), obfuscating parameters via gauge randomization, reducing black-box attack success, and preserving predictive accuracy close to unprotected models.
-
Interpretability extensions from the TT representation. The TT form allows efficient computation of marginal and conditional distributions, supporting feature-sensitivity analysis, construction of cancer-type-specific models without retraining, and interpretability for otherwise black-box models such as NNs.
Main Findings
-
LORIS leaks its training cohort. WB attack scores on released LORIS parameters identify Cho1 with score 1.0000, and coefficients reconstructed from the public web interface identify Cho1 with 0.9944 and Cho2 with 0.8138 (Cho1 and Cho2 are train/test partitions of the same original dataset, so high scores for both are expected).
-
Cross-validation amplifies leakage rather than reducing it. Averaged LRs trained via repeated cross-validation score higher than vanilla LRs at every access level (bBB 0.9149 vs. 0.8178; cBB 0.9910 vs. 0.9129; WB 0.9999 vs. 0.9330), despite similar AUC values.
-
Deeper access means more leakage. Unprotected vanilla LR attack Hamming scores rise from 0.8178 (bBB) to 0.9129 (cBB) to 0.9330 (WB). NN WB attacks score 0.6336, lower than the same model's bBB (0.7375) and cBB (0.8608) scores, which the authors attribute to difficulty extracting structured information from more complex parameter spaces.
-
Tensorization drives white-box attacks to chance. Direct WB attacks on gauge-randomized TT parameters yield accuracies close to 50%: TT-LR scores are 0.5117, 0.5112, and 0.5104 for b = 2, 6, 10; TT-NN scores are 0.5061 (b = 2), 0.5018 (b = 6), and 0.5025 (b = 10).
-
Tensorization gives black-box protection comparable to practical DP baselines while retaining accuracy. At b = 2, TT-LR yields bBB 0.6666 and cBB 0.8231 with AUCs of 0.68/0.70/0.67/0.59/0.61/0.62 across the six cohorts, comparable to LR-DP with ε ∈ (1, 10) but with higher AUC than those DP settings (which range from 0.50 to 0.75).
-
Larger b and larger ε both increase leakage. For TT-LR, bBB rises from 0.6666 (b = 2) to 0.7535 (b = 6) to 0.7687 (b = 10). For LR-DP, bBB rises from 0.5314 (ε = 0.1) to 0.5710 (ε = 1) to 0.7163 (ε = 10) to 0.7663 (ε = 100).
-
Strong DP guarantees cost substantial utility. LR-DP at ε = 0.1 has AUCs of 0.51, 0.51, 0.50, 0.50, 0.50, 0.50 across cohorts, versus 0.74–0.76 for unprotected vanilla LR. Among the tested DP configurations, ε ≈ 10 offers the best empirical privacy–utility balance.
-
NN-DP utility loss is partly pipeline-driven. Predictive performance remains hindered even at ε = ∞ (AUCs 0.74/0.74/0.67/0.64/0.62/0.72), which the authors attribute to differences in the DP-SGD pipeline, such as the reduced number of epochs or gradient clipping, rather than noise addition.
-
A cohort of 35 patients can be detected. In a task distinguishing models trained on Cho1 (964 samples) from Cho1 plus the Kato cohort (35 patients), averaged LRs reach 0.9182 under cBB and 1.0000 under WB on the Kato label; vanilla LRs reach 0.7141 under WB. TT-LR at b = 2 remains at 0.5375 (bBB), 0.5811 (cBB), and 0.4966 (WB), and TT-NN at b = 2 stays near chance at 0.4779, 0.5226, and 0.5152.
-
LR-specific coefficient reconstruction is a residual risk for TT-LR. Because LR scores are monotonic, LR coefficients can be reconstructed from TT evaluations, improving with larger b: starred WB scores are 0.7461 (b = 2), 0.7979 (b = 6), and 0.8129 (b = 10). These attacks do not outperform black-box attacks on the original LR, and this strategy requires the adversary to know the underlying architecture and is not readily extensible to NNs.
-
Interpretability results are only partially available. The paper states that TT approximations preserve key properties of LORIS such as response monotonicity and enable marginals, conditionals, feature-sensitivity analysis, and cancer-type-specific models without retraining, but the numerical interpretability results in Section 4 are not included in the material available for this summary.
Methodology in Plain English
The authors assume an adversary who knows the model architecture and training procedure (with some hyperparameter uncertainty), has access to six public patient cohorts, and can train as many shadow models as needed.
-
Attack construction. For each hyperparameter configuration and each possible combination of cohorts, they train 100 shadow models. Each shadow model is paired with a membership label: a binary vector whose m-th entry is 1 if cohort C_m was in that model's training data. An adversarial MLP multi-label classifier (three hidden layers of sizes 32, 16, 8; output layer of size 6, one per public cohort) learns to predict that vector from model information alone.
-
Three access levels. Under bBB the adversary sees only binary classifications; under cBB, continuous output probabilities; under WB, full model parameters. For BB attacks, every shadow model is evaluated on the same 100 samples drawn from the union of all cohorts. For WB attacks the adversary receives 22 parameters for LR (21 coefficients plus intercept), an 818-dimensional concatenated vector for the NN, and concatenated TT cores of 168 dimensions (TT-LR) or 1,020 dimensions (TT-NN).
-
Target models. LRs are trained as vanilla (single 80% split) or averaged (20 repetitions of 3-fold cross-validation, coefficients averaged across folds, matching the LORIS procedure), and NNs are 2-layer MLPs with two hidden layers of size 19.
-
Defense. For each trained model, the authors use TT-RSS to build a tensor train from 50 random training-set pivots (80 pivots and ranks r = 5 for NNs). Model evaluations on the pivots are discretized into b bins (b ∈ {2, 6, 10}), with values below 0.5 mapped to the lower bin limit and values above 0.5 to the upper limit, preserving the property that output probabilities sum to 1. The resulting TTs have N = 22 cores, ranks r = 2, input dimension d = 2, and polynomial embeddings φ(x) = [1, x]. TT cores are then randomized by a gauge transformation, so white-box access reveals no more than black-box behavior. TT-RSS costs O(|D|²Nd) model evaluations on pivot dataset D plus O(|D|³Nd) to assemble the TT.
-
Comparison baselines. LR-DP models are trained from scratch with solver "lbfgs" and penalty "l2" at ε ∈ {0.1, 1, 10, 100}; NN-DP uses DP-SGD with max_grad_norm = 1, δ = 10⁻⁴, and σ ∈ {20, 5, 1, 0} (approximately ε ∈ {0.2, 1, 10, ∞}), with epochs reduced to 50.
-
Evaluation. Attack success is reported as the Hamming score (proportion of correctly predicted cohort-membership labels; 0.5 = random guessing, 1.0 = perfect identification), averaged over five repetitions. Clinical performance is reported as median balanced accuracy (using Youden's J threshold) and AUC. All experiments ran on an Intel Xeon CPU E5-2620 v4 with 256 GB RAM and an NVIDIA GeForce RTX 3090, using Scikit-Learn, Diffprivlib, PyTorch, Opacus, and TensorKrowch.
Why This Matters
Impact on research. The paper reframes model transparency as a privacy liability in clinical ML: it shows that the cross-validation averaging procedures commonly recommended for generalization can make models more identifiable, and it positions tensorization as a post-hoc, architecture-agnostic alternative to differential privacy that acts as a form of knowledge distillation. It also connects formal white-box privacy guarantees for tensor networks to a concrete clinical use case.
Real-world applications:
-
Publicly hosted clinical risk calculators and web interfaces. The demonstrated attack against LORIS via its public web interface shows that releasing or exposing logistic regression parameters can reveal which patient cohorts were used, even from rounded probability outputs.
-
Rare-disease and small-cohort studies. Because a 35-patient cohort such as Kato can be reliably detected within datasets of hundreds to thousands, hospitals contributing small or rare cancer subtype datasets face identifiable disclosure risk.
-
Deployment of interpretable models in regulated clinical settings. Tensorized models retain monotonicity, feature importance, and support marginal/conditional analysis, which matters where regulators and clinicians require explainable predictions.
-
Interpretability for black-box models. The same tensorization pipeline works on NNs that cannot be recovered from their TT parameters, giving an interpretability path for models that would otherwise be opaque.
Industry relevance. Organizations hosting clinical prediction services, health systems contributing patient cohorts to multi-institutional studies, and developers of privacy-preserving ML tooling all have a stake: the paper argues that privacy and interpretability need not be traded off, and provides an open-source implementation for reproducibility.
Future Directions
-
Theoretical bounds linking b to recoverable information. The authors suggest deriving formal bounds on the information recoverable for a given discretization value b, analogous to the role of ε in differential privacy.
-
Output obfuscation with formal DP guarantees before tensorization. Since tensorization is independent of the obfuscation method, the authors propose replacing discretization with noise-based output obfuscation that carries DP guarantees, which they hypothesize could provide stronger protection, particularly against coefficient reconstruction in TT-LR.
-
Scaling and stability of TT-RSS. Tensorization occasionally produces degenerate models with accuracies near 50 percent, attributed to the intrinsic randomness of TT-RSS and mitigable by increasing the pivot dataset size, which raises computational cost; the paper flags this tradeoff as open.
-
Extending attacks to unknown cohorts and higher-dimensional settings. The attack assumes membership among a known set of candidate cohorts rather than discovery of arbitrary unknown cohorts, and the authors note that the disparate accuracy impact of DP noise may become more pronounced in higher-dimensional settings.
Target Audience
Researchers and practitioners in clinical machine learning, privacy-preserving ML, and medical AI deployment; privacy and security engineers evaluating membership inference risk in released models; methodologists working on tensor network models and quantum-inspired machine learning; and clinicians, regulators, or data stewards who need to weigh the transparency of models like logistic regression against the privacy of the patient cohorts used to train them.
Authors’ abstract
Publicly available clinical machine learning models pose an underappreciated privacy risk: their parameters or outputs can be exploited to recover information from patients whose data were used during training. Moreover, this risk is exacerbated by models such as logistic regression (LR), which are typically preferred in clinical settings for their transparency. To assess this empirically, we attack LORIS, a publicly available LR model for immunotherapy response prediction hosted on a U.S. government website. From evaluations through its public interface, we recover the model parameters and identify the training cohort with high confidence. More broadly, we design cohort-level membership inference attacks under three levels of adversarial access---binary black-box, continuous black-box, and white-box---and apply them to both LR models and shallow neural networks (NNs) trained on the same task. Our results reveal that even a cohort of 35 patients can be reliably identified within training sets of hundreds to thousands, and that common practices such as cross-validation amplify rather than mitigate this risk. To address these vulnerabilities, we propose a quantum-inspired defense based on tensorizing discretized models into tensor trains (TTs). This representation obfuscates model parameters and preserves accuracy, while offering black-box protection comparable to practical Differential Privacy baselines. Additionally, the TT representations retain LR interpretability and extend it through efficient computation of marginal and conditional distributions, enabling this richer analysis also for black-box models such as NNs. Our results establish tensorization as a practical, post-hoc tool for private, interpretable, and effective clinical prediction.