Research
RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection
Overview Research area: Computer vision, specifically AI-generated video detection, digital watermarking, and media forensics. Technical level: Intermediate. The paper is readable for a general machin

- arXiv
- 2512.10248
- Published
- 2025-12-11
- Authors
- Zhuo Wang, Xiliang Liu, Ligang Sun
AI summary
Overview
- Research area: Computer vision, specifically AI-generated video detection, digital watermarking, and media forensics.
- Technical level: Intermediate. The paper is readable for a general machine-learning audience, but it assumes familiarity with detection metrics (true positive rate, true negative rate, F1), transformer video backbones, multimodal large language models (MLLMs), and paired statistical testing.
- Scope: The paper introduces RobustSora, a 6,500-video benchmark built to measure how much AI-generated video detectors rely on embedded provenance watermarks rather than on genuine generation artifacts.
What This Paper Is About
Commercial video generators such as Sora 2 embed visible or semi-transparent overlay watermarks in their output for provenance tracking, and detectors trained on those outputs may simply learn "watermark present means AI-generated." The authors build a controlled benchmark that removes watermarks from generated videos and injects fake watermarks into authentic videos, then measure how much detection accuracy shifts. The goal is to determine, with causal evidence rather than conjecture, whether state-of-the-art detectors depend on watermark cues and whether that dependency can be reduced through training.
Key Contributions
- The first watermark robustness benchmark for AIGC video detection. RobustSora is presented as the first benchmark explicitly designed to evaluate how watermark presence and absence affect detector performance.
- A systematically constructed four-type dataset of 6,500 manually verified videos. The categories are Authentic-Clean (A-C), Generated-Watermarked (G-W), Generated-DeWatermarked (G-DeW), and Authentic-Spoofed (A-S), drawn from authentic sources (Vript, DVF, UltraVideo) and generators (Sora, Sora 2, Pika, Open-Sora 2, KLing).
- A dual-task evaluation protocol. Task-I (Watermark Erasure Robustness) tests detection on watermark-removed AI videos; Task-II (Watermark Spoofing Robustness) measures false alarms on authentic videos carrying fake watermarks.
- Cross-architecture analysis with causal controls and a mitigation. Ten models spanning specialized detectors, transformer classifiers, and MLLM-based approaches are benchmarked, backed by a placebo control that bounds inpainting-artifact confounds at ≤ 2 pp and a watermark-aware training augmentation that recovers 3–4 pp on both tasks.
Main Findings
- Baseline ranking: On the Standard Test (A-C vs. G-W with watermarks intact), DuB3D-FF achieves the highest overall accuracy (0.824) and F1 (0.823), followed by MViT V2 (0.801). The 2.3 pp gap between them is statistically significant (paired bootstrap p < 0.01) with standard deviations ≤ 0.010 for both.
- Watermark erasure hurts detection (Task-I): Removing watermarks from AI videos reduces AI-detection accuracy for every model, with per-model drops ranging from 3.7 pp (Qwen2.5-VL-3B) to 9.4 pp (Video-LLaVA-7B). DuB3D-FF drops 9.1 pp, MViT V2 drops 8.1 pp, and TimeSformer drops 7.7 pp.
- Watermark spoofing causes false alarms (Task-II): Faking watermarks on authentic videos degrades authentic-detection accuracy for most models, from 2.6 pp (Qwen2.5-VL-7B, not significant) to 9.0 pp (VideoSwin-T). DeCoF drops 7.2 pp, D3 drops 7.5 pp, and MViT V2 drops 8.2 pp.
- One counterintuitive exception: Qwen2.5-VL-3B improves by 1.6 pp on Task-II (reported as not significant), which the figure caption describes as a paradoxical watermark interaction.
- Aggregate effect size: Across the ten models, watermark manipulation induces accuracy changes of −9.4 to +1.6 pp, with a mean of 6.6 pp, and p < 0.01 for 7 of 10 models on each task.
- Confounds are bounded: A placebo condition in which DiffuEraser is applied to a non-watermark patch of equal area bounds inpainting-artifact effects at ≤ 2 pp, well below the 5–9 pp drop caused by genuine watermark removal.
- Training intervention works: A watermark-aware training augmentation recovers 3–4 pp on both tasks.
- Generator prominence, not architecture, drives dependency: Per-generator breakdown shows Sora 2 induces drops of −11 to −14 pp versus −3 to −6 pp for Pika and Open-Sora 2.
- Spoofing generalizes across templates: Under a second, KLing-derived watermark template, detectors' vulnerability rankings correlate with the Sora 2 template at Spearman ρ = 0.94.
- Architecture-specific profiles: Specialized detectors cluster at 0.72–0.76 overall accuracy, with NSG-VD showing the most balanced profile (0.719 AI accuracy / 0.774 authentic accuracy); transformer models range from 0.733 (VideoSwin-T) to 0.824 (DuB3D-FF); MLLMs lag, with Qwen2.5-VL-3B at 0.261 AI accuracy versus 0.834 authentic accuracy and Video-LLaVA-7B showing the reverse (0.712 AI vs. 0.602 authentic).
- Baseline reproduction is sound: Each baseline passes a sanity check on a 500-video GenVideo subset, recovering 87–95% accuracy in agreement with the original publications, so the lower absolute numbers on RobustSora are attributed to a harder distribution-shift setting rather than reproduction error.
Methodology in Plain English
The authors collected 3,000 authentic camera-captured videos and 3,500 AI-generated videos (6,500 total), standardizing everything to 3–10 seconds and 480p–1080p. Authentic footage came from Vript (1,200 YouTube plus 300 TikTok), DVF (800), and UltraVideo (700). Generated footage came from Sora 2 (1,070, which carries an overlay watermark), Pika (520 text-to-video plus 280 image-to-video), Open-Sora 2 (600), KLing (530), and Sora (500). All authentic videos were manually verified to be free of synthetic artifacts and watermarks.
To make the de-watermarked set, the team localized watermarks by computing a temporal-median residual, flagged candidate pixels above a threshold set at the 95th percentile of that residual on a held-out clip, discarded connected components smaller than 100 pixels or larger than 10% of the frame, dilated the region by 4 pixels, and applied a 3-frame temporal max-pool to build a stable mask. DiffuEraser then inpainted the masked regions; ProPainter served as a second tool for the artifact-control experiment. Every video was manually inspected, and the 12.4% that failed were reprocessed with relaxed thresholds.
To make the spoofed set, they extracted a watermark template from Sora 2 clips by taking per-frame residuals and robust medians across frames and clips, jointly fitting an alpha blending strength and spatial offset, and alpha-blending that template onto all 3,000 authentic videos at a constant location per clip, with intensity jittered between 0.9 and 1.1 times the fitted alpha. A KLing template served as an independent check.
Splitting was designed to mimic deployment: 2,400 A-C plus 2,800 G-W videos (5,200 total) for training, with no exposure to manipulated variants. Testing used three conditions — Standard Test (600 A-C + 700 G-W), Task-I (700 G-DeW), and Task-II (600 A-S). Ten detectors were trained under identical conditions: specialized AIGC detectors (DeCoF, NSG-VD, D3), transformer classifiers (TimeSformer, VideoSwin-T, MViT V2, DuB3D-FF), and MLLM approaches (Qwen2.5-VL-3B/7B, Video-LLaVA-7B). Non-MLLM models used AdamW at a learning rate of 1×10⁻⁴, batch size 32, and 8 frames for spatial models or 16 for spatio-temporal models. MLLMs used LoRA fine-tuning with rank 32 and scaling factor 64, a chain-of-thought prompt, 16 frames, AdamW at 2×10⁻⁵, and a cosine schedule. Results are reported as means over three random seeds, with differences from the Standard Test tested via paired bootstrap over n = 10,000 resamples.
Why This Matters
Impact on research. The paper reframes a benchmark artifact as a causal confound. Prior AIGC video benchmarks such as GenVidBench (143,000 videos, 79.90% with MViT V2), AEGIS (over 10,000 videos, Qwen2.5-VL at 22–23% zero-shot on the hardest subsets), and DuB3D (2.66 million videos, 96.77% in-domain and 79.19% out-of-domain) intermix watermarked and non-watermarked clips without accounting for the difference. RobustSora positions itself against this blind spot and against VideoMarkBench, which measures whether watermarks survive perturbation; RobustSora instead measures whether detectors lean on them, describing the combination as a doubly vulnerable paradigm where a fragile signal supports a dependent classifier.
Real-world applications:
- Content moderation and platform integrity pipelines that flag synthetic video, where a watermark-removal step by an adversary could silently degrade detection.
- Provenance and authenticity verification tools, since spoofed watermarks on real footage can trigger false positives and erode trust in legitimate content.
- Journalism and fact-checking workflows that need calibrated confidence in "AI-generated" verdicts rather than watermark presence.
- Forensic reporting and legal evidence handling, where a detector's stated accuracy must hold when provenance markings have been stripped or forged.
Industry relevance. Because the dependency is strongest for the most prominent watermark (Sora 2 induced drops of −11 to −14 pp versus −3 to −6 pp for Pika and Open-Sora 2), generator providers and detector vendors share the risk. The finding that a watermark-aware training augmentation recovers 3–4 pp gives vendors a concrete, low-cost mitigation. The authors state that the dataset, evaluation code, removal and spoofing pipelines, and pretrained checkpoints will be publicly released.
Future Directions
- Hardening detection training. The 3–4 pp recovery from watermark-aware augmentation suggests a broader question: what training recipe removes watermark dependency entirely without sacrificing clean-video accuracy.
- Cross-template and cross-generator generalization. The KLing template check (Spearman ρ = 0.94) is one step; whether spoofing vulnerability rankings hold across additional watermark styles, resolutions, and compression pipelines is untested here.
- Closing the MLLM gap. Qwen2.5-VL-3B at 0.261 AI-detection accuracy and Video-LLaVA-7B's inverted profile show that parameter-efficient tuning of general-purpose multimodal models leaves forensic capability underexplored, including whether watermark sensitivity scales with model capacity.
- Rethinking watermarking under adversarial pressure. Since removing a watermark degrades detection and forging one causes false alarms, an open question is whether any watermarking scheme can supply a signal robust enough to be worth depending on, given the failure modes documented in VideoMarkBench.
Target Audience
This paper is most useful to researchers building and evaluating AI-generated content detectors, benchmark designers in media forensics, and engineers deploying deepfake or synthetic-media detection in production. It also speaks to watermarking researchers who need to understand the downstream consumption of their signals, to policy and standards groups working on content provenance, and to graduate students entering the AIGC detection field who need a clear demonstration of shortcut learning in forensic classifiers.
Authors’ abstract
The proliferation of AI-generated video models poses new challenges to information integrity and digital trust. A key confound, however, remains unaddressed: commercial generators embed visible overlay watermarks for provenance tracking, yet no existing benchmark controls for this variable, leaving open whether detectors learn genuine generation artefacts or merely associate watermark patterns with AI-generated labels. We present RobustSora, a benchmark of 6,500 manually verified videos in four categories: Authentic-Clean (A-C), Generated-Watermarked (G-W), Generated-DeWatermarked (G-DeW), and Authentic-Spoofed (A-S), sourced from Vript, DVF, and UltraVideo (authentic) and from Sora, Sora 2, Pika, Open-Sora 2, and KLing (generated). Two evaluation tasks isolate watermark effects: Task-I (Watermark Erasure Robustness) tests detection on watermark-removed AI videos; Task-II (Watermark Spoofing Robustness) measures false-alarm rates on authentic videos injected with fake watermarks. Across ten models spanning specialized detectors, transformer classifiers, and MLLMs, watermark manipulation induces accuracy changes of $-9.4$ to $+1.6$ pp (mean 6.6 pp; $p{<}0.01$ for 7/10 models on each task). A placebo control bounds inpainting-artefact confounds at $\le$2 pp, and a watermark-aware training augmentation recovers 3-4 pp on both tasks, together providing causal evidence that detectors actively rely on watermark cues. Per-generator breakdown shows that Sora 2 induces drops of $-11$ to $-14$ pp versus $-3$ to $-6$ pp for Pika and Open-Sora 2, indicating that watermark prominence, rather than detector architecture, is the principal driver of dependency. These results argue for watermark-aware evaluation and training in AIGC video detection. Dataset, evaluation code, and pretrained checkpoints will be released.