Research
Robustness Certificates for Neural Networks Against Data Poisoning and Evasion Attacks
Overview Research area: Formal robustness certification for machine learning — specifically, using control-theoretic safety verification (barrier certificates) to certify neural networks against data

- arXiv
- 2512.20865
- Published
- 2025-12-24
- Authors
- Sara Taheri, Mahalakshmi Sabanayagam, Debarghya Ghoshdastidar, Majid Zamani
AI summary
Overview
Research area: Formal robustness certification for machine learning — specifically, using control-theoretic safety verification (barrier certificates) to certify neural networks against data poisoning at training time and evasion attacks at test time.
Technical level: Advanced. The paper builds on discrete-time dynamical systems, barrier certificates from control theory, scenario convex programs, and probably approximately correct (PAC) bounds.
Scope: The paper proposes a model-agnostic, data-driven framework that learns a neural network-based barrier certificate (NNBC) to compute a certified ℓ_p-norm poisoning radius, guaranteeing that test-accuracy degradation stays below a prescribed threshold α, with a PAC-style probabilistic guarantee.
What This Paper Is About
Most defenses against data poisoning are heuristic and lack formal guarantees, while existing certification methods typically assume a fixed corruption budget, target specific model classes (e.g., decision trees, nearest neighbors, graph neural networks), require white-box access, or certify only individual test points. This paper asks whether one can certify, for any ML model, an ℓ_p-norm bounded poisoning budget such that performance degradation stays below a threshold α. The authors answer yes by modeling gradient-based training as a discrete-time dynamical system and recasting poisoning robustness as a formal safety verification problem solved with barrier certificates.
Key Contributions
-
Training as a dynamical system. The authors cast gradient-based ML training as a discrete-time dynamical system (dt-DS), where model parameters form the system state and the (possibly poisoned) training data act as the input, and reformulate robustness certification against train-time and test-time perturbations as a formal safety verification problem using barrier certificates (BCs).
-
Neural network-based barrier certificate (NNBC). Because explicit BC design is intractable in high-dimensional parameter spaces with unknown poisoned training dynamics, the authors parameterize the BC as a neural network (NNBC) trained on finite sets of poisoned trajectories. The NNBC yields the certified robust radius — the largest admissible perturbation of train or test data for which test-accuracy degradation is provably at most a given threshold.
-
PAC guarantee via scenario convex program. Verification is reformulated as a scenario convex program (SCP), producing a probably approximately correct (PAC) bound that gives a rigorous probabilistic guarantee on the trained NNBC and its associated certified robust radius.
-
Model-agnostic and unified. The framework requires no prior knowledge of the ML architecture, the attack strategy, or the amount of data corrupted, and it extends to test-time (evasion) certification — claimed in the abstract to be the first unified framework providing formal guarantees in both training-time and test-time attack settings.
Main Findings
-
Certified robust radius exists under mild conditions: Theorem 12 states that if a BC satisfying the conditions in Definition 9 exists for perturbations with ‖Δ‖_p ≤ δ_cert (resp. ‖Δ′‖_p ≤ δ′_cert), then δ_cert (resp. δ′_cert) is a certified train-time (resp. test-time) robust radius, guaranteeing that under worst-case perturbations the accuracy degradation is at most α.
-
Barrier conditions are threefold: The BC must satisfy (10) ℬ(θ) ≤ 0 for all θ in the initial set Θ_0, (11) ℬ(θ) > 0 for all θ in the unsafe set, and (12) the forward-invariance condition ℬ(f(θ, Δ)) ≤ 0 for all θ with ℬ(θ) ≤ 0. The paper notes condition (12) is a drift/separation-style discrete-time barrier condition, deliberately not a contraction-type condition, since contraction can be overly restrictive for data-driven training dynamics.
-
A necessary bound from sampling: The empirical radius δ_emp is the largest perturbation bound for which all sampled terminal parameters satisfy 𝒢(θ(t_∞)) ≤ α. Because it comes from finite sampling, δ_emp is an optimistic upper bound on true robustness, so any valid certificate must satisfy δ_cert ≤ δ_emp.
-
Certification is control-theoretic, unlike prior work: The related-work review groups existing poisoning certificates into four categories — ensemble-based methods, randomized smoothing for poisoning, differential-privacy-based methods, and model-specific certification methods — and notes that these are predominantly not control-theoretic, whereas this framework is explicitly barrier-certificate-based verification of training dynamics.
-
Empirical claim from the abstract: Experiments on MNIST, SVHN, CIFAR-10, and CIFAR-100 show the approach certifies non-trivial perturbation budgets while being model-agnostic and requiring no prior knowledge of the attack or contamination level. The specific certified radii, accuracy numbers, and experimental tables are not reported in the provided text of this paper (the content is truncated in Section III-C).
-
Threat model coverage: Train-time poisoning perturbs an unknown fraction ρ ∈ [0,1] of training samples (r := ⌈ρn⌉ inputs modified), while test-time evasion perturbs an unknown fraction ρ′ ∈ [0,1] of test samples (r′ := ⌈ρ′n′⌉ inputs modified) with model parameters held fixed. Both attacks are bounded in ℓ_p norm by δ and δ′ respectively. The framework focuses on input-space poisoning (features, not labels), and although the exposition assumes the first r training and r′ test elements are poisoned, the framework is permutation-invariant and does not require knowledge of which indices are corrupted.
-
Safety is defined in state space: The parameter space is partitioned into a safe set Θ_s^{Δ′} := {θ ∈ ℝ^d | 𝒢(θ) ≤ α} and its complement, the unsafe set. The primary objective is terminal-time performance — ensuring θ(t_∞) satisfies the safety condition — although the definition is stated at the level of parameters rather than a specific iteration.
Methodology in Plain English
The authors first reframe training. Instead of viewing training as a procedure that produces a model, they view it as a dynamical system: the parameters θ are the state, each gradient update step is the state transition, and the training data is the input. Because the data may be poisoned, the input can be adversarially perturbed.
They then borrow a tool from control theory — the barrier certificate. A barrier certificate is a function that is non-positive on all "good" starting points, strictly positive on all "bad" parameter configurations, and stays non-positive along every possible training trajectory. If such a function exists, no poisoned training run can ever push the parameters into the unsafe region. The largest perturbation budget for which this holds becomes the certified robust radius.
Since the true training dynamics and poisoning attack are unknown, the barrier certificate cannot be derived analytically. So the authors learn it: they train the model many times under N uniformly spaced poisoning budgets (one grid for train-time perturbations, one for test-time, with the other fixed to zero), record the initial parameters, the terminal parameters, and whether each terminal model was safe or unsafe relative to the threshold α. A neural network — the NNBC, with ReLU activations in hidden layers and identity activation in the output layer — is then trained on this data using a composite loss, for a candidate radius δ_cand no larger than the empirical radius δ_emp.
Finally, because the NNBC only sees a finite sample of trajectories, the authors reformulate the verification conditions as a scenario convex program and derive a PAC bound. This bound says that with confidence at least 1 − β, the probability that the barrier conditions are violated on unseen trajectories is at most ε — turning a finite-data fit into a probabilistic guarantee. The same machinery applies to test-time evasion, since the certificate concerns parameter trajectories only and is agnostic to model architecture, optimizer, attack strategy, and attack strength.
Why This Matters
Impact on research. The work connects control-theoretic safety verification to adversarial machine learning, offering a certification route that does not depend on model class, attack type, known corruption level, or point-wise reasoning. It also targets a gap the authors identify: formal certificates for data poisoning are far less developed than those for test-time adversarial attacks. The abstract positions the framework as the first unified approach giving formal guarantees in both training-time and test-time attack settings.
Real-world applications:
- Autonomous driving, where poisoned training data can embed stealth vulnerabilities that trigger failures in safety-critical operation.
- Medical diagnostics, cited as another safety-critical deployment area where training data flows through third-party datasets and labeling workflows.
- Supply-chain data pipelines, where datasets are sourced, refreshed, or re-annotated by third parties under a bounded corruption budget and can be influenced upstream.
- Periodic retraining on newly collected data, where corruption can persist unnoticed and degrade downstream model performance.
Industry relevance. Practitioners deploying ML in regulated or safety-critical settings need quantified guarantees rather than empirical attack benchmarks. A method that returns a certified perturbation budget — the largest attack magnitude for which accuracy degradation is provably bounded by α — is directly actionable for risk assessment and for setting data-validation budgets, and the model-agnostic, no-knowledge-of-attack property reduces the effort needed to apply it to existing gradient-trained pipelines.
Future Directions
- Closing the gap between certified and empirical radii. The paper establishes only the necessary condition δ_cert ≤ δ_emp, leaving open how tight certified radii are relative to the sampled optimistic upper bound in practice.
- Extending beyond input-space poisoning. The framework currently focuses on perturbations to input features; label poisoning, backdoor triggers, and joint feature-and-label corruption are handled by other method families in the related work and are not covered here (the detailed loss and verification construction is truncated in the provided content, so the full scope of what is certified is not fully visible).
- Scaling and tightening the PAC bound. Since the guarantee comes from a scenario convex program over a finite sample, the sample size N and the resulting ε and β trade-off are natural levers to study — how the certified radius shrinks as confidence requirements tighten is not reported in the available text.
- Broadening the model and optimizer classes. The framework is stated for models trained with gradient-based optimizers, and future work would logically test how far this extends across architectures, optimizers, and stopping criteria (the terminal horizon t_∞ is defined by epoch count or a stopping criterion).
Target Audience
Researchers and graduate students in trustworthy machine learning, adversarial robustness, and formal verification; control theorists interested in data-driven safety certificates for learning systems; and practitioners in safety-critical domains (autonomous systems, medical ML) who need formal, quantified robustness guarantees rather than empirical defenses. Readers should be comfortable with dynamical systems notation, ℓ_p norms, convex optimization, and PAC-style generalization arguments.
Authors’ abstract
The increasing use of machine learning in safety-critical domains amplifies the risk of adversarial threats, especially data poisoning attacks that corrupt training data to degrade performance or induce unsafe behavior. Most existing defenses lack formal guarantees or rely on restrictive assumptions about the model class, attack type, extent of poisoning, or point-wise certification, limiting their practical reliability. This paper introduces a principled formal robustness certification framework that models gradient-based training as a discrete-time dynamical system (dt-DS) and formulates poisoning robustness as a formal safety verification problem. By adapting the concept of barrier certificates (BCs) from control theory, we introduce sufficient conditions to certify a robust radius ensuring that the terminal model remains safe under worst-case ${\ell}_p$-norm-based poisoning. To make this practical, we parameterize BCs as neural networks trained on finite sets of poisoned trajectories. We further derive probably approximately correct (PAC) bounds by solving a scenario convex program (SCP), which yields a confidence lower bound on the certified robustness radius generalizing beyond the training set. Importantly, our framework also extends to certification against test-time attacks, making it the first unified framework to provide formal guarantees in both training and test-time attack settings. Experiments on MNIST, SVHN, CIFAR-10, and CIFAR-100 show that our approach certifies non-trivial perturbation budgets while being model-agnostic and requiring no prior knowledge of the attack or contamination level.