Research
Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
Overview Research area: Computer vision, specifically face forgery (deepfake) detection using parameter-efficient fine-tuning of the DINOv2 vision foundation model. Technical level: Advanced. The pape

- arXiv
- 2511.12107
- Published
- 2025-11-15
- Authors
- Tianxiang Zhang, Peipeng Yu, Zhihua Xia, Longchen Dai, Xiaoyu Zhou, Hui Gao
AI summary
Overview
- Research area: Computer vision, specifically face forgery (deepfake) detection using parameter-efficient fine-tuning of the DINOv2 vision foundation model.
- Technical level: Advanced. The paper assumes familiarity with Vision Transformers, LoRA-style low-rank adaptation, mixture-of-experts routing, and cross-dataset evaluation protocols.
- Scope: This paper proposes the DeepFake Fine-Grained Adapter (DFF-Adapter), a 3.5M-parameter tuning scheme that injects multi-head LoRA adapters into every Transformer block of a frozen DINOv2 backbone and jointly trains authenticity detection with forgery-type classification.
What This Paper Is About
Most deepfake detectors built on large pre-trained vision models treat forgery detection as a single binary classification problem, ignoring the fact that different manipulation methods leave different traces. The authors argue that this discards useful information about which kind of forgery was used, and that existing fine-tuning approaches also restrict adaptation to the last Transformer block or a few final layers, which limits how much task-specific signal reaches early layers. Their goal is a lightweight fine-tuning method that learns fine-grained, method-specific artifact cues and uses them to make the authenticity decision more generalizable to unseen forgeries.
Key Contributions
- A DeepFake Fine-Grained Adapter (DFF-Adapter) architecture that intertwines a shared branch with task-specific branches, so the authenticity detector inherits fine-grained forgery cues from the auxiliary forgery-type classification task. The adapter is inserted into query, value, and dense projections of every Transformer block of a frozen DINOv2 backbone.
- A Forgery-Aware Multi-Head Router (FAMHR, also referred to in the text as DF-MHR) that partitions intermediate Transformer features into multiple channel subspaces and, for each subspace, dynamically routes to a learned top-3 set of LoRA experts drawn from a shared pool. This per-subspace expert allocation is intended to mine localized forgery artifacts and enable multi-view feature fusion.
- A Shared-Enhanced Task Fusion (SETF) module that residually fuses shared and task-specific updates across all blocks, transferring fine-grained cues from the auxiliary branch to the authenticity stream.
- Empirical results showing strong generalization with only 3.5M trainable parameters: state-of-the-art-average performance across cross-dataset benchmarks and the best overall performance among compared methods on the DF40 cross-manipulation benchmark.
Main Findings
- Intra-dataset accuracy: On FF++ (trained on the c23 compression version), the method reaches a detection AUC of 99.56, described by the authors as highly competitive under the standard DeepfakeBench evaluation protocol.
- Cross-dataset generalization: Trained on FF++ and tested on unseen datasets, the method records AUCs of 95.26 (CDF-v2), 89.96 (DFDC), 96.14 (CDF-v1), and 91.57 (DFDCP), giving an average of 93.23 and what the paper describes as an average improvement of 2.41 AUC points over compared state-of-the-art methods. Among the baselines, LVLM-DFD scores 99.53 / 94.71 / 79.12 / 97.62 / 91.81, averaging 90.82.
- Cross-manipulation on DF40: Evaluated across three forgery categories covering 12 manipulation methods, the method leads all compared approaches on every reported method. Selected AUCs: 93.15 (FaceDancer), 93.98 (InSwapper), 98.84 (FSGAN), 94.51 (UniFace), 95.36 (FOMM), 90.28 (HyperReenact), 91.65 (Wav2Lip), 86.29 (MCNet), 99.92 (StyleGAN3), 100.00 (StyleGAN-XL), 100.00 (VQGAN), 98.44 (DiT-XL/2). For comparison, SBI scores 78.18 / 88.52 / 89.62 / 89.02 / 88.03 / 65.26 / 77.06 / 81.51 / 97.91 / 23.26 / 91.47 / 53.59, and LVLM-DFD scores 82.97 / 87.64 / 93.75 / 90.61 / 93.34 / 81.56 / 78.60 / 83.45 / 98.87 / 100.00 / 99.99 / 86.61.
- Ablation confirms both modules matter: A vanilla DINOv2 fine-tuned setup yields 61.93 (CDF-v2), 52.83 (DFDC), 68.46 (CDF-v1), 55.54 (DFDCP). Adding FAMHR raises these to 89.56 / 86.24 / 85.53 / 87.11, and adding SETF on top gives the full 95.26 / 89.96 / 96.14 / 91.57.
- Outperforms other fine-tuning paradigms: Against Linear Probing (66.63 / 72.54 / 56.45 / 65.90), plain LoRA (85.14 / 79.95 / 85.14 / 81.49), and MoE-FFD (80.40 / 76.52 / 65.75 / 87.60) on CDF-v2 / DFDC / CDF-v1 / DFDCP, the full method again reports 95.26 / 89.96 / 96.14 / 91.57.
- Works with very few identities: Using only 10 randomly sampled FF++ training identities (about 1% of available identities), AUCs are 81.97 (CDF-v2), 82.63 (DFDC), 81.62 (CDF-v1), 82.22 (DFDCP); 30 identities (about 3%) give 81.94 / 79.16 / 90.49 / 81.25; 50 identities (about 5%) give 87.28 / 82.37 / 91.26 / 86.16.
- Highest AUC and lowest EER against pre-trained model baselines: With DFF-Adapter, DINOv2 ViT-L/14 reports FF++ 99.56 AUC / 1.79 EER, CDF-v2 95.26 / 12.65, DFDC 89.96 / 17.96. CLIP ViT-B/16 with the adapter reports 99.29 / 0.71, 85.19 / 22.47, 83.19 / 24.82, and CLIP ViT-L/14 with the adapter reports 97.14 / 2.57, 89.10 / 15.28, 84.75 / 23.32. DINOv2 ViT-B/16 with the adapter reports 99.61 / 2.32, 93.35 / 14.61, 82.00 / 25.49.
- Hyperparameter choices were validated by ablation. For loss weights (lambda_0 = 10), setting lambda_1 = 2 gave the best CDF-v2 AUC (95.26) and DFDC AUC (89.96) among the tested settings of 0, 2, 10, and 50. For head adapter count, h = 4 gave the best CDF-v2 and DFDC AUCs among h = 1, 4, 8. For LoRA experts per head, 6 gave the best CDF-v2 and DFDC AUCs among 1, 4, 6, and 12.
- Feature separability: t-SNE visualizations on FF++ and CDF-v2 show that a DINOv2 model fine-tuned only at the last layer produces entangled real/fake clusters, whereas the DFF-Adapter produces compact, well-separated clusters that also distinguish forgery types.
- Interpretability: Grad-CAM visualizations on CDF2, DFDC, and DFDCP indicate the method focuses on manipulated local regions of faces more consistently than the DINOv2 baseline.
- Entire face synthesis results: On a set of DF40 GAN- and diffusion-based synthesis methods, the method reports AUCs of 92.31 (DiT-XL/2), 98.44 (RDDM), 100.00 (PixArt-alpha), 94.65 (SiT-XL/2), 99.92 (StyleGAN3), 100.00 (StyleGAN-XL), and 100.00 (VQGAN). The text describes selecting eight representative methods, but seven are named and shown in the table.
- Stated limitation: The method assumes different manipulation methods produce distinct and consistent artifact patterns; in real-world settings artifacts may overlap, be subtle, or evolve with newer generators, potentially adding noise to the auxiliary task and weakening the main detection stream.
Methodology in Plain English
The authors start from a frozen DINOv2 ViT-L/14 backbone and, instead of fine-tuning it, attach small adapter modules inside every Transformer block. Each adapter contains three low-rank (LoRA) "heads": one for binary authenticity detection, one for forgery-type classification, and one shared head that participates in both tasks. Rather than using a single LoRA per head, the head's input channels are split into several subspaces, and a router scores a pool of low-rank experts and activates only the top three for each subspace, then combines their outputs. The shared head's output is concatenated with the task-specific outputs to form the classification tokens, creating a residual path that carries fine-grained forgery-type cues into the authenticity decision. Training uses two forward passes per mini-batch over the same images — one with the task flag set to authenticity, one with it set to forgery type — so the branch gradients do not interfere, and the two losses are combined with weights lambda_0 = 10 and lambda_1 = 2. Only the adapters and task-specific classifiers are updated; the backbone stays frozen. Configuration details: rank r = 16, scaling factor alpha = 32, 6 LoRA experts, 4 head adapters, images cropped to 224 x 224, batch size 24, 50 epochs, a single NVIDIA RTX 4090 GPU, Adam optimizer with learning rate 2 x 10^-4 and weight decay 1 x 10^-5. Evaluation uses video-level AUC, with EER also reported in some tables; results were trained only on the c23 version of FF++.
Why This Matters
Research impact. The paper argues that deepfake detection with foundation models has been under-specified as a binary task, and shows that adding an auxiliary manipulation-type objective with a shared transfer path improves cross-dataset and cross-manipulation generalization while using only 3.5M trainable parameters. It also suggests that adapting all Transformer blocks, rather than only the final one, matters for catching low-level forgery cues.
Real-world applications.
- Social media and news platforms screening uploaded video for synthetic manipulation before distribution.
- Identity verification and remote onboarding systems that need to reject spoofed or reenacted faces.
- Legal and journalistic forensics, where graded evidence about which manipulation technique was used can support an investigation.
- Privacy-preserving biometric pipelines, where training data is scarce and few-identity training is the practical constraint — the paper's 10-identity experiments show usable performance in that regime.
Industry relevance. The low trainable-parameter count (3.5M) and single-GPU training setup make this approach practical for organizations that cannot afford full fine-tuning of large vision backbones. The finding that performance holds up with as few as 10 to 50 training identities is directly relevant to sectors facing tightening restrictions on collecting face data, and the cross-manipulation results target the reality that new generative methods appear faster than detectors can be retrained.
Future Directions
- Handling newer and overlapping manipulation types. The authors state they will enhance detection performance on data generated by the latest models, and their own limitation section identifies overlapping, subtle, or evolving artifact patterns as a threat to the forgery-type branch.
- Real-world deployment algorithms. The paper lists designing efficient algorithms to tackle real-world scenarios as future work, which would require testing beyond curated benchmarks such as DeepfakeBench and DF40.
- Resolving ambiguity in manipulation-type labels. A natural extension is to make the auxiliary classification task robust when artifacts are shared across methods or intentionally mimic authentic distributions, since the current design assumes distinct and consistent patterns.
- Broadening the architecture comparison. The paper reports results for DINOv2 ViT-B/16 and ViT-L/14 and for CLIP with and without the adapter; whether the design transfers to other frozen backbones or scales to additional public benchmarks is not reported.
Target Audience
Researchers and engineers working on media forensics, deepfake detection, and parameter-efficient adaptation of vision foundation models. It is also relevant to practitioners in platform trust-and-safety, biometric verification, and digital provenance who need a detection model that generalizes to unseen generators under limited training data and compute. Readers without background in Vision Transformers, LoRA, or mixture-of-experts routing will find the methods section demanding, since the paper presents routing equations, gating scores, and top-3 expert selection without introductory scaffolding.
Authors’ abstract
The proliferation of sophisticated deepfakes poses significant threats to information integrity. While DINOv2 shows promise for detection, existing fine-tuning approaches treat it as generic binary classification, overlooking distinct artifacts inherent to different deepfake methods. To address this, we propose a DeepFake Fine-Grained Adapter (DFF-Adapter) for DINOv2. Our method incorporates lightweight multi-head LoRA modules into every transformer block, enabling efficient backbone adaptation. DFF-Adapter simultaneously addresses authenticity detection and fine-grained manipulation type classification, where classifying forgery methods enhances artifact sensitivity. We introduce a shared branch propagating fine-grained manipulation cues to the authenticity head. This enables multi-task cooperative optimization, explicitly enhancing authenticity discrimination with manipulation-specific knowledge. Utilizing only 3.5M trainable parameters, our parameter-efficient approach achieves detection accuracy comparable to or even surpassing that of current complex state-of-the-art methods.