Skip to content
AI.info

Research

Optimizing DINOv2 with Registers for Face Anti-Spoofing

Overview Research area: Face anti-spoofing (presentation attack detection), specifically detecting physical (print, replay, cutout, mask) and digital (face-swap, attribute editing, GAN-generated) atta

Optimizing DINOv2 with Registers for Face Anti-Spoofing
arXiv
2510.17201
Published
2025-10-20
Authors
Mika Feng, Pierre Gallin-Martel, Koichi Ito, Takafumi Aoki

AI summary

Overview

  • Research area: Face anti-spoofing (presentation attack detection), specifically detecting physical (print, replay, cutout, mask) and digital (face-swap, attribute editing, GAN-generated) attacks before a face recognition system authenticates a user.
  • Technical level: Advanced. The work assumes familiarity with Vision Transformers, self-supervised pretraining (DINOv2), attention mechanics, and anti-spoofing evaluation protocols.
  • Scope in one sentence: The paper proposes a DINOv2 ViT-B/14 backbone augmented with four register tokens and fine-tuned through only its last encoder block to detect minute live-versus-spoof differences, evaluated on the 6th Face Anti-Spoofing Workshop dataset and the SiW dataset.

What This Paper Is About

Face recognition systems tolerate natural variation in pose, illumination, and blur, which also makes them vulnerable when an attacker shows a photo or replay of a registered user. The paper's goal is to detect these spoofing attempts by capturing the very small texture, pattern, and structural differences between a live face and a fake one. The authors argue that general-purpose Vision Transformer features, when cleaned up with register tokens and only lightly fine-tuned, generalize better to spoofing attack types the model never saw during training than conventional CNN- and ViT-based detectors do.

Key Contributions

  1. A face anti-spoofing method built on DINOv2 ViT-B/14 with registers, using the pretrained model to extract generalizable features and the register tokens to suppress attention perturbations so the network focuses on essential, minute cues.
  2. A partial unfreezing strategy: only the last of the 12 encoder blocks is fine-tuned, which the authors report preserves generalization and produces the best test-set ACER in their ablation (0.1107 versus 0.1443 and 0.2147 for unfreezing more blocks).
  3. A training recipe for the anti-spoofing task: separate learning rates for the classification head (5×10⁻⁵) and the DINOv2 backbone (5×10⁻⁶), AdamW, cosine annealing, and focal loss with gamma = 2 and equal class weights to handle strong class imbalance.
  4. An empirical demonstration on two datasets — the FAS Workshop dataset (15 attack categories) and SiW (Protocols 1, 2, and 3) — showing improved ACER, AUC, and ACC against the workshop baseline and improved ACER against eight conventional methods on the hardest SiW protocol.

Main Findings

  • FAS Workshop dataset ablation: Reducing the number of trainable layers lowered accuracy on the validation dataset but raised accuracy on the test dataset. Validation ACER rose from 0.2552 (unfreezing layers 10–12) to 0.2875 (layers 11–12) to 0.3159 (layer 12 only), while test ACER fell from 0.2147 to 0.1443 to 0.1107 across the same settings. The authors attribute this to a significant distribution shift between the validation and test sets.
  • Baseline comparison (FAS Workshop dataset): The proposed method beat the baseline (AjianLiu) on ACER (0.1107 versus 0.2259), AUC (0.9480 versus 0.8989), and ACC (0.9047 versus 0.6355), but the baseline was better on EER (0.2307 versus 0.6132).
  • Explanation for the EER gap: The authors state that the proposed method's poor EER stems from an imbalance between APCER and BPCER. Because EER is the point where the false acceptance rate equals the false rejection rate, a large gap between these two metrics forces the threshold into a position that is suboptimal for each.
  • SiW Protocol 3 (unknown attacks): The lowest ACER among conventional methods was 0.83 ± 0.13%, while the proposed method reached 0.625 ± 0.495% (APCER 0.625 ± 0.565%, BPCER 0.585 ± 0.455%).
  • SiW Protocols 1 and 2: All compared methods showed low error values, indicating these are easy protocols. The proposed method reported 0.00 APCER, BPCER, and ACER on Protocol 1, and 0.00 ± 0.00 on all three metrics for Protocol 2.
  • Practical implication of the register tokens: The paper reports that a "spike phenomenon" in ViT attention — abnormally high attention on non-essential patches such as background — is caused by outlier tokens with extremely high vector norms that carry global image information but lack local features. Register tokens act as dedicated storage for this temporary information, removing the need to reuse unnecessary patches as memory.

Methodology in Plain English

The team takes an off-the-shelf DINOv2 ViT-B/14 model that has been pretrained with register tokens and adapts it to anti-spoofing. An input image is resized to 224×224 pixels, split into 14×14 patches, and passed through a 12-layer encoder. A classification head is attached after the last layer, trained with a two-class loss on the class token, and at inference the head's output gives the live probability.

Registers matter because Vision Transformers tend to dump global, non-local information into a few outlier tokens, which distorts the attention map and pulls focus toward background patches. Four learnable register tokens give the model somewhere to park that information, so attention stays on the fine textures and patterns that distinguish live from spoofed faces.

For training, the authors fine-tune with AdamW using a larger learning rate on the classification head and a smaller one on the backbone (5×10⁻⁵ and 5×10⁻⁶), apply cosine annealing, and use focal loss with gamma = 2 and equal class weights. On the FAS Workshop dataset they run 200 epochs with early stopping at patience 20 and a mini-batch of 32, freezing everything except the last encoder block. On SiW they change the setup: the entire model is unfrozen, the loss becomes cross entropy, the optimizer becomes Nesterov Stochastic Gradient Descent, and augmentation combines a traditional method with FAS-Aug applied at probability 0.2. For SiW, five frames are randomly extracted from each training or evaluation video per protocol, faces are cropped using the known bounding boxes, and images are resized to 224×224.

Why This Matters

  • Impact on research: The paper tests whether large self-supervised vision backbones designed for general visual understanding can be repurposed for the very fine-grained forensic task of spoof detection, and it isolates which design choices (register tokens, minimal fine-tuning) preserve the cross-attack generalization that CNN-based detectors like CDCN++, NAS-FAS, and PatchNet lose on unseen attacks.
  • Real-world applications:
    • Mobile and banking face-unlock or identity-verification flows that must reject printed photos, screen replays, and cutout masks.
    • Remote identity onboarding (KYC), where attackers increasingly submit deepfake or face-swapped imagery generated digitally.
    • Access control at physical checkpoints such as airports or secure facilities using 3D masks or transparent/plaster/resin presentation attacks.
    • Forensic or moderation pipelines that screen user-uploaded face media for manipulation.
  • Industry relevance: Anti-spoofing sits directly in front of commercial face recognition. A lower ACER on unknown attacks (0.625 ± 0.495% versus 0.83 ± 0.13% for the best conventional method on SiW Protocol 3) is the metric vendors care about most, because production systems face attack types that were never in the training distribution. The finding that an entire DINOv2 backbone can be adapted by training only one of its 12 blocks also matters for cost, since it reduces the compute needed to adapt large pretrained models to a new domain.

Future Directions

  • Fixing the APCER/BPCER imbalance: The proposed method's EER of 0.6132 against the baseline's 0.2307 on the FAS Workshop dataset shows the two error rates are far apart. Calibrating or reweighting them could remove the one metric on which the method loses.
  • Closing the validation-to-test gap: The ablation shows validation accuracy and test accuracy moving in opposite directions as more layers are unfrozen, which the authors call a distribution shift. Better validation splits or domain-adaptation strategies could make model selection more reliable.
  • Reconciling the training recipes: The FAS Workshop experiments fine-tune one block with focal loss and AdamW, while SiW unfreezes everything with cross entropy and Nesterov SGD. Understanding whether a single recipe serves both is left open.
  • Reducing the variance on SiW Protocol 3: The proposed method's 0.625 ± 0.495% ACER carries high variance relative to the Feng method's 0.83 ± 0.13%, and the reported APCER (0.625 ± 0.565%) exceeds its BPCER (0.585 ± 0.455%). Stabilizing results across frames and seeds is a natural next step.

Target Audience

Researchers and engineers working on presentation attack detection, biometric security, and deepfake or digital-forgery detection, particularly those already using Vision Transformer backbones or foundation models in their pipelines. It is also useful for practitioners who need to adapt a large pretrained model to a small, heavily imbalanced dataset without losing generalization, and for graduate students following the line of work from CDCN++ and TransFAS through to foundation-model-based anti-spoofing.

Authors’ abstract

Face recognition systems are designed to be robust against variations in head pose, illumination, and image blur during capture. However, malicious actors can exploit these systems by presenting a face photo of a registered user, potentially bypassing the authentication process. Such spoofing attacks must be detected prior to face recognition. In this paper, we propose a DINOv2-based spoofing attack detection method to discern minute differences between live and spoofed face images. Specifically, we employ DINOv2 with registers to extract generalizable features and to suppress perturbations in the attention mechanism, which enables focused attention on essential and minute features. We demonstrate the effectiveness of the proposed method through experiments conducted on the dataset provided by ``The 6th Face Anti-Spoofing Workshop: Unified Physical-Digital Attacks Detection@ICCV2025'' and SiW dataset. The project page is available at: https://gsisaoki.github.io/FAS-DINOv2-ICCVW/ .

Read the original paper