Skip to content
AI.info

Research

BLANKET: Anonymizing Faces in Infant Video Recordings

Overview Research area: Computer vision, specifically privacy-preserving video anonymization (de-identification) for a demographic group (infants) underrepresented in existing datasets and models. Tec

BLANKET: Anonymizing Faces in Infant Video Recordings
arXiv
2512.15542
Published
2025-12-17
Authors
Ditmar Hadera, Jan Cech, Miroslav Purkrabek, Matej Hoffmann

AI summary

Overview

Research area: Computer vision, specifically privacy-preserving video anonymization (de-identification) for a demographic group (infants) underrepresented in existing datasets and models.

Technical level: Intermediate. The paper combines diffusion-model inpainting, landmark detection, and face swapping with downstream evaluation using pose estimation, but each component is explained without deep mathematics.

One-sentence scope: The paper proposes BLANKET, a two-stage pipeline that replaces an infant's identity in video with a newly generated, compatible synthetic identity while preserving expression, gaze, head pose, and other facial attributes, and evaluates it against DeepPrivacy2 on a dataset of 46 infant videos.

What This Paper Is About

Sharing video of human subjects is needed for developmental research, but it requires anonymization for ethical and privacy reasons. Trivial approaches such as blurring or placing a black box over the face remove all facial information (expression, gaze, age, gender, race) and can also harm downstream analysis such as pose estimation. This paper builds an anonymization method tuned specifically for infants, whose facial proportions and features differ from adults and are poorly handled by existing generative models and face recognition engines.

Key Contributions

  1. A new anonymization method, BLANKET (Baby-face Landmark-preserving ANonymization with Keypoint dEtection consisTency), tuned for infants. It is designed to seamlessly change identity, preserve facial expressions and other attributes, produce minimal perceptual artifacts, and remain temporally consistent across video frames.
  2. A two-stage pipeline in which a new compatible random identity is generated by inpainting the first detected face with Stable Diffusion, and that identity is then swapped into every frame using FaceFusion to obtain temporally consistent, expression-preserving results.
  3. An extensive quantitative and human evaluation of the method against DeepPrivacy2, covering de-identification, preservation of gender/race/emotion/gaze/eye and mouth openness/head orientation, video-consistency statistics, and a user study with about 40 participants.
  4. An assessment of impact on a downstream task, human pose estimation, using RTMDet-l for person detection and ViTPose-b for pose estimation, plus a black-rectangle baseline. The code is released at https://github.com/ctu-vras/blanket-infant-face-anonym.

Main Findings

  • De-identification (automatic metric): ArcFace cosine distance between original and anonymized frames is 0.11 ± 0.18 for BLANKET and 0.19 ± 0.26 for DeepPrivacy2. The authors state the difference is not very significant and hypothesize ArcFace does not generalize well to infants and newborns, so they rely on the user study and qualitative results for this aspect.
  • Attribute preservation: BLANKET scores higher on same race ratio (0.69 ± 0.19 versus 0.56 ± 0.18) and same emotion ratio (0.51 ± 0.13 versus 0.27 ± 0.11); same gender ratio is close (0.81 ± 0.21 versus 0.79 ± 0.16). Gender, race, and emotions are estimated with the DeepFace library.
  • Gaze and openness: BLANKET shows smaller gaze difference (0.36 ± 0.37 versus 0.39 ± 0.38), smaller eye openness difference (0.094 ± 0.098 versus 0.13 ± 0.12), and smaller mouth openness difference (0.11 ± 0.15 versus 0.17 ± 0.18).
  • Head orientation: Angular differences are lower for BLANKET on all three axes: X 0.14 ± 0.32 versus 0.45 ± 0.57 rad, Y 0.08 ± 0.11 versus 0.18 ± 0.16 rad, Z 0.08 ± 0.13 versus 0.17 ± 0.15 rad.
  • Video consistency: Original identity variance is 0.016 ± 0.031 for both. Anonymized identity variance is 0.025 ± 0.039 for BLANKET versus 0.051 ± 0.063 for DeepPrivacy2. Landmark trajectory correlation with the original is 0.956 ± 0.064 for BLANKET versus 0.86 ± 0.14 for DeepPrivacy2.
  • Person detection: Detection AP is 98.1 with no anonymization, 50.9 with a black rectangle, 81.5 with DeepPrivacy2, and 90.7 with BLANKET. The authors note 90 AP is very high compared with the state of the art on COCO (66.0 AP).
  • Pose estimation: Pose AP is 100.0 (no anonymization), 18.1 (black rectangle), 79.1 (DeepPrivacy2), 97.2 (BLANKET). Excluding the five facial keypoints: 100.0, 38.4, 92.3, 97.7 respectively. In-the-wild AP (detection plus pose on anonymized images): 100.0, 17.3, 72.3, 91.7. Detection contributes more to the performance loss than pose estimation; publishing ground-truth bounding boxes with anonymized faces would preserve nearly 98% of pose estimation performance.
  • User study, perceived de-identification: DeepPrivacy2 averages 4.5 on a 1–5 scale (1 = same identity, 5 = completely different); BLANKET averages 2.5. The authors suggest the small resolution of the footage partly explains this, since only the inner face is changed.
  • User study, artifacts: Users judged BLANKET as producing fewer artifacts than the competing method in 96% of cases. Original frames were identified as original in 96% of cases, BLANKET frames were mistaken for original in 24% of cases, and DeepPrivacy2 frames were never confused with the original.
  • Qualitative behavior: BLANKET handles occlusion and off-plane head rotations and mostly preserves expression; the authors observed that occasionally the model does not generate fully closed eyes. DeepPrivacy2 suffers severe spatial and temporal artifacts, changes the face dramatically between frames, and the authors conclude it is unusable for infants.

Methodology in Plain English

The input is an original video plus a target image representing a new identity. Stage one builds that target identity. The pipeline takes the first frame in which a face is detected, finds all faces with YOLO11, locates 98 facial landmarks and a head pose with the SPIGA model using "wflw" weights, builds a binary mask as the convex hull of those landmarks, and inpaints the masked face region using Stable Diffusion. Because the inpainting has little or no access to the region inside the mask, it synthesizes a random identity that fits the surrounding head and context. Two checkpoints are used: Realistic Vision V2.0 and Realistic Vision V6.0 B1. A positive prompt "a face of a baby" and a standard negative prompt are used, along with CFG and ControlNet to fit the landmarks and face geometry; parameters were tuned empirically to trade off de-identification against compatibility. A high noise level tends to produce visible artifacts and a too-low level tends to produce identities too close to the original.

Stage two inserts that identity into every frame of the video with FaceFusion, which detects, tracks, and aligns faces and performs the swap, including lip-syncing, expression-matching, and a face enhancement model. Inpainting takes about 8 seconds on a consumer GPU but runs only once; FaceFusion runs close to real time.

For evaluation, the authors use a subset of the Infant Pose Estimation dataset by Chambers et al., presented in a later work: 46 videos of infants, each approximately 100 seconds, totaling 78 minutes and about 140k frames. For the pose-estimation experiments they sample frames uniformly from the 46 videos, resulting in 8897 images; detecting persons with RTMDet-l and estimating pose for boxes with confidence above 0.3 using ViTPose-b, followed by pose-based non-maximum suppression, gives 9102 pseudo ground-truth annotations. AP follows the COCO evaluation protocol and is computed globally, so no standard deviation is reported for those tables. The user study used two questionnaires, each containing 5 videos and 9 images, filled by about 40 users, with randomly chosen (not cherry-picked) media and shuffled presentation order.

Why This Matters

The paper addresses a practical bottleneck in developmental psychology and behavioral research: spontaneous daily recordings of infants are needed to study development, but privacy and ethics concerns make sharing raw footage difficult, and existing anonymization tools were not built for infant faces. The work also shows that the choice of anonymization method materially affects downstream computer vision results, and that the dominant source of failure is person detection rather than pose estimation itself.

Real-world applications:

  • Sharing infant and child video datasets for developmental psychology and behavior analysis while preserving usable facial and bodily information.
  • Privacy compliance for clinical or home-monitoring recordings of babies, where blurring or black boxes would destroy clinically relevant facial expression and gaze cues.
  • Human pose and behavior analysis pipelines that need to run on anonymized footage without large accuracy loss, including detection and keypoint estimation.
  • Benchmarking anonymization tools for non-adult subjects, since the paper documents how a method that works on adults (DeepPrivacy2) fails on infants.

Industry relevance: the released code and the reported near-real-time FaceFusion stage make the approach relevant to teams building privacy-preserving video ingestion for datasets, monitoring products, or research archives. The paper's finding that publishing ground-truth bounding boxes alongside anonymized faces preserves nearly 98% of pose estimation performance is directly actionable for dataset release practices.

Future Directions

  1. Handling face detection failures. Anonymization relies on accurate face detection; if detection fails or another face is detected in the frame, the target may remain unmodified. Currently such cases are handled by obscuring the entire frame; future work will explore interpolation from adjacent anonymized frames.
  2. Stronger de-identification. Users occasionally perceived BLANKET's anonymization as insufficient because the face resembles the original, partly due to low resolution and because the method only alters the facial interior, retaining landmarks that carry identity cues.
  3. Automated selection of dissimilar but compatible identities. The authors propose exploring metrics for choosing new identities that are maximally dissimilar to the original yet still compatible, noting that face recognition engines perform poorly on infants, which makes this hard to automate.
  4. Detecting and correcting perceptual artifacts, following prior work, to further reduce artifacts introduced by the inpainting stage.

Target Audience

Researchers and practitioners in computer vision and privacy who work on face anonymization, generative models, or video processing; developmental psychologists and behavioral scientists who need to share video recordings of infants ethically; engineers building data release or monitoring pipelines that must comply with privacy requirements; and anyone evaluating anonymization methods for non-adult or otherwise underrepresented subjects, for whom off-the-shelf tools such as DeepPrivacy2 may not generalize.

Authors’ abstract

Ensuring the ethical use of video data involving human subjects, particularly infants, requires robust anonymization methods. We propose BLANKET (Baby-face Landmark-preserving ANonymization with Keypoint dEtection consisTency), a novel approach designed to anonymize infant faces in video recordings while preserving essential facial attributes. Our method comprises two stages. First, a new random face, compatible with the original identity, is generated via inpainting using a diffusion model. Second, the new identity is seamlessly incorporated into each video frame through temporally consistent face swapping with authentic expression transfer. The method is evaluated on a dataset of short video recordings of babies and is compared to the popular anonymization method, DeepPrivacy2. Key metrics assessed include the level of de-identification, preservation of facial attributes, impact on human pose estimation (as an example of a downstream task), and presence of artifacts. Both methods alter the identity, and our method outperforms DeepPrivacy2 in all other respects. The code is available as an easy-to-use anonymization demo at https://github.com/ctu-vras/blanket-infant-face-anonym.

Read the original paper