Skip to content
AI.info

Research

TopoReformer: Mitigating Adversarial Attacks Using Topological Purification in OCR Models

Overview Research area: Adversarial machine learning and computer vision, specifically adversarial defense for Optical Character Recognition (OCR) systems using topological data analysis. Technical le

TopoReformer: Mitigating Adversarial Attacks Using Topological Purification in OCR Models
arXiv
2511.15807
Published
2025-11-19
Authors
Bhagyesh Kumar, A S Aravinthakashan, Akshat Satyanarayan, Ishaan Gakhar, Ujjwal Verma

AI summary

Overview

Research area: Adversarial machine learning and computer vision, specifically adversarial defense for Optical Character Recognition (OCR) systems using topological data analysis.

Technical level: Advanced. The paper assumes familiarity with adversarial attack formulations (FGSM, PGD, C&W, EOT, BPDA), persistent homology, Vietoris–Rips complexes, and autoencoder training objectives.

One-sentence scope: The paper introduces TopoReformer, a model-agnostic purification pipeline built on a topological autoencoder that removes adversarial perturbations from text images before they reach an OCR classifier, without adversarial retraining.

What This Paper Is About

OCR systems are vulnerable to adversarial perturbations that are invisible to humans but cause wrong transcriptions, and some of these perturbations survive real-world print-and-scan or display-and-camera capture. Existing defenses — adversarial training, input preprocessing, and post-recognition correction — tend to be model-specific, computationally expensive, degrade performance on clean inputs, and fail against adaptive attacks. The goal of this work is a drop-in purification front end that suppresses adversarial artifacts while keeping the structural integrity of text images, trained only on unperturbed data.

Key Contributions

  1. Topological purification. A Topological Autoencoder (TopoAE) purifies adversarial images by enforcing consistency between the persistent homology of the input space and its latent representation, followed by a lightweight Reformer (a VAE) that re-aligns the purified output to the manifold the classifier expects.
  2. Freeze-Flow training paradigm. A training scheme that blocks direct gradients to the primary encoder and routes them through an auxiliary module, forcing reliance on topology-consistent latents. The paper reports up to 5% improvement in classification performance under Carlini–Wagner attacks from this component.
  3. Model-agnostic OCR defense. A drop-in pipeline for any OCR model, integrating the topology-guided purifier and Reformer, that provides robustness against white-box, black-box, and adaptive attacks without requiring adversarial training and with classifier weights kept fixed.
  4. First use of topological autoencoders for this purpose. The authors state that, to the best of their knowledge, this is the first work to employ topological autoencoders for end-to-end adversarial purification or defense, noting that prior topological work addressed logit alignment or other defense formulations.

Main Findings

  • Classical attacks (MNIST, EMNIST). Under weak/strong Carlini–Wagner (c = 1e-2 / 1e+1), an undefended classifier reaches F1 of 30.41% / 4.30% on MNIST and 36.54% / 33.85% on EMNIST. Adding the full pipeline (+Warmup) raises this to 65.86% / 75.15% on MNIST and 69.66% / 68.82% on EMNIST, with matching precision gains (73.17% / 77.31% and 73.20% / 72.51%).
  • Incremental ablation. Each component adds measurable benefit on Carlini attacks: MNIST F1 moves from 53.92%/48.51% (+TopoAE) to 65.38%/67.93% (+Reformer) to 65.36%/72.41% (+Aux) to 65.86%/75.15% (+Warmup) for the weak/strong settings.
  • PGD and FGSM saturate. At ε = 0.005 / 0.01, undefended PGD F1 is already 96.74% / 96.62% on MNIST and 84.87% / 72.66% on EMNIST, so absolute gains are small (e.g., 97.70% / 97.62% and 84.53% / 83.79% with the full pipeline). The paper attributes this to iterative PGD and multi-step FGSM often saturating on standard datasets.
  • Anomalies in the ablation. EMNIST PGD and FGSM results with +Aux and +Warmup (e.g., PGD F1 83.25% / 82.75%) fall below +Reformer (90.30% / 89.51%). The paper attributes such anomalies to manifold mismatch between the TopoReformer output and the classifier's inputs.
  • Adaptive attacks. Under EOT, the undefended model shows ASR 99.05% (MNIST) and 97.73% (EMNIST) with F1 of 3.38% and 1.42%; TopoReformer reduces ASR to 9.19% and 28.32% with F1 of 90.73% and 73.28%. Under EOT+BPDA, ASR falls from 99.70% and 98.69% to 36.59% and 44.26%. Under BPDA alone, the defense weakens substantially: ASR 81.14% on MNIST and 84.46% on EMNIST, with F1 of 15.65% and 12.77%.
  • Clean-input performance preserved. The paper reports approximately 98% accuracy on MNIST and 94% on EMNIST in unperturbed settings, arguing that robustness is gained without sacrificing standard classification capability — unlike most adversarial training approaches.
  • OCR-specific FAWA attack. Baseline OCR models show near-total vulnerability (ASR 100% for CRNN, 99.83% Rosetta, 98.92% STAR-Net, 99.92% RARE, 99.83% TRBA). With TopoReformer, ASR drops to 78.83% (CRNN), 44.08% (Rosetta), 65.17% (STAR-Net), 87.67% (RARE), and 60.75% (TRBA), with character accuracy rising to 71.00%, 85.98%, 79.81%, 65.18%, and 80.26% respectively. The Reformer component was deliberately excluded from OCR evaluations to preserve deployability efficiency.
  • Latent topology visualization. Even when trained with moderate noise additions, the TopoReformer produces consistent, separable latent representations. MNIST clusters are described as clean and radially separable, while EMNIST clusters appear more entangled due to higher class diversity but remain discernible.
  • Grad-CAM interpretability. Under adversarial noise, baseline attention is scattered or shifted; TopoReformer keeps compact, semantically consistent activation. In the example shown, an adversarial input predicted "j" with confidence 0.18 while the correct label was "i"; the reformed output predicts "i" with confidence 0.88. The paper also notes that confidence scores of unperturbed images increase after processing through the TopoAE.

Methodology in Plain English

The pipeline has three parts. First, a Topological Autoencoder takes the input image and produces both a latent vector and a reconstructed, purified image. During its training, it computes persistence diagrams — summaries of how connected components, loops, and voids appear and disappear as a scale parameter grows — for both the input point cloud and its latent representation, and penalizes any mismatch between them. Because adversarial perturbations tend to distort only local pixel relationships while leaving global structure intact, this penalty makes the model discard changes that do not correspond to meaningful topological variation.

Second, a Reformer (a VAE) takes the purified image and adjusts it so it lands closer to the manifold the downstream classifier expects, using a combined objective of mean squared error against the TopoAE output, cross-entropy against ground-truth labels from the classifier's logits, and a KL divergence regularizer. Third, an Auxiliary Module receives the topological latent vector and injects it into the bottleneck of the Reformer, giving the reformer access to topology-aware information that is similar for clean and perturbed inputs.

A key training trick is Freeze-Flow: when the auxiliary pathway is added naively, gradients do not propagate to it well because the model relies on the purified-image path. So the encoder of the Reformer VAE is frozen for a warmup period to force gradients through the auxiliary module, after which the decoder is unfrozen and the whole VAE plus auxiliary module are trained jointly. The TopoAE is trained first on unperturbed data alone until convergence, then frozen; the classifier weights never change.

For OCR specifically, word-level images lacked consistent geometric structure, so words were decomposed into characters, auto-padded to a uniform 2×28×44 size, purified at the character level, then recombined into word-level format of 32×100.

Evaluation used MNIST, EMNIST letters, and an OCR dataset spanning 53 classes (uppercase, lowercase, and a blank token), generated with open-source code. MNIST and EMNIST classifiers follow the MAGNET architectures; OCR models span three CTC-based designs (CRNN, Rosetta, STAR-Net) and two attention-based designs (RARE, TRBA). Training used the Adam optimizer with a learning rate of 0.001 and loss weights of 1, 0.5, and 0.5.

Why This Matters

Impact on research. The paper argues that robustness can emerge from preserving global manifold topology rather than from explicit gradient regularization, and that this implicitly constrains Lipschitz continuity — limiting how abruptly latent representations vary under input perturbations. It also positions topology-based regularization as a complementary axis of robustness (manifold-level rather than gradient-level) and provides an evaluation that deliberately includes adaptive attacks to avoid the false sense of security produced by gradient obfuscation.

Real-world applications.

  • Document automation and processing pipelines, where transcription errors can propagate to financial loss or policy breaches.
  • License plate recognition and traffic enforcement, where errors can lead to wrongful enforcement actions.
  • Automated compliance systems that depend on correct extraction from scanned or photographed documents.
  • Enterprise data extraction from documents captured through print-and-scan or display-and-camera channels, where adversarial perturbations can survive capture.

Industry relevance. Because TopoReformer is model-agnostic, requires no adversarial retraining, and leaves classifier weights fixed, it can act as a drop-in front end for deployed OCR stacks — including both CTC-based and attention-based architectures. The paper reports that clean-input performance is not sacrificed, which addresses a major practical objection to adversarial defenses in production systems.

Future Directions

  • Integrating topology-aware training directly into OCR model optimization, removing the need for a separate Reformer stage.
  • Constructing a comprehensive benchmark dataset and standardized classifier suite for OCR-specific adversarial attacks and defenses, covering both classical and adaptive threat models, to address the gap in publicly available resources.
  • Closing the BPDA gap: the paper hypothesizes that the topological regularizer enforces global Lipschitz-like smoothness but does not directly constrain the fine-grained local curvature that BPDA exploits, which leaves an open problem for the defense design.
  • Broadening evaluation beyond MNIST, EMNIST, and the single 53-class OCR dataset, and extending the OCR evaluation to include the Reformer component, which was excluded to preserve efficiency.

Target Audience

Researchers and practitioners in adversarial machine learning, robust computer vision, and document AI; engineers deploying OCR in security- or compliance-sensitive pipelines; and readers with a background in topological data analysis who are interested in how persistent homology can be turned into a practical defense mechanism. The paper is not beginner-friendly: understanding the loss formulations, persistence pairings, and adaptive attack protocols requires prior exposure to both adversarial ML and computational topology.

Authors’ abstract

Adversarially perturbed images of text can cause sophisticated OCR systems to produce misleading or incorrect transcriptions from seemingly invisible changes to humans. Some of these perturbations even survive physical capture, posing security risks to high-stakes applications such as document processing, license plate recognition, and automated compliance systems. Existing defenses, such as adversarial training, input preprocessing, or post-recognition correction, are often model-specific, computationally expensive, and affect performance on unperturbed inputs while remaining vulnerable to unseen or adaptive attacks. To address these challenges, TopoReformer is introduced, a model-agnostic reformation pipeline that mitigates adversarial perturbations while preserving the structural integrity of text images. Topology studies properties of shapes and spaces that remain unchanged under continuous deformations, focusing on global structures such as connectivity, holes, and loops rather than exact distance. Leveraging these topological features, TopoReformer employs a topological autoencoder to enforce manifold-level consistency in latent space and improve robustness without explicit gradient regularization. The proposed method is benchmarked on EMNIST, MNIST, against standard adversarial attacks (FGSM, PGD, Carlini-Wagner), adaptive attacks (EOT, BDPA), and an OCR-specific watermark attack (FAWA).

Read the original paper