Research
HP-GAN: Harnessing pretrained networks for GAN improvement with FakeTwins and discriminator consistency
Overview Research area: Generative adversarial networks (GANs) for image synthesis, specifically computer vision generative modeling that combines pretrained feature networks, self-supervised learning

- arXiv
- 2602.03039
- Published
- 2026-02-03
- Authors
- Geonhui Son, Jeong Ryong Lee, Dosik Hwang
AI summary
Overview
- Research area: Generative adversarial networks (GANs) for image synthesis, specifically computer vision generative modeling that combines pretrained feature networks, self-supervised learning (SSL), and multi-discriminator training.
- Technical level: Advanced. The paper assumes familiarity with GAN min-max training, Barlow Twins-style SSL, hinge losses, spectral normalization, and CNN/ViT feature backbones, although its two core ideas are conceptually simple.
- Scope: The paper introduces HP-GAN, a FastGAN-based framework that adds a self-supervised "FakeTwins" loss applied to generated images and a "discriminator consistency" regularizer between CNN-based and ViT-based discriminators, and evaluates it on seventeen datasets spanning large, small, few-shot, and AFHQ regimes using FID, KID, precision, recall, and perceptual path length.
What This Paper Is About
GANs are hard to train: they suffer from non-convergence, instability, and mode collapse, and these problems get worse when data or compute is scarce (the paper cites medical imaging of rare diseases, specific celebrity portraits, and a particular artist's artwork). Recent work uses pretrained networks for perceptual losses or as frozen feature spaces, but the authors argue these networks are used narrowly. HP-GAN's goal is to use pretrained networks in a second, additional role as SSL encoders that directly train the generator (FakeTwins), and to make multiple heterogeneous discriminators agree with one another (discriminator consistency), producing more diverse and higher-quality generated images.
Key Contributions
- Discriminator consistency loss: A new regularization term that reduces disparities in the outputs of discriminators operating on CNN-derived and ViT-derived feature maps, giving the generator more uniform and coherent feedback and stabilizing training (Config D, then Config E).
- FakeTwins: A self-supervised mechanism that uses fixed pretrained feature networks as SSL encoders and applies a Barlow Twins-style loss through generated (fake) images to train the generator directly, with the aim of increasing diversity and fidelity.
- Incremental architecture design: A progression from FastGAN (Config A) to Projected GAN-style multi-discriminator training with a CNN feature network (Config B), to a hybrid CNN + ViT feature-network setup with small latent code and blur regularization (Config C), then discriminator consistency (Config D), then FakeTwins (HP-GAN, Config E).
- Extensive evaluation: Experiments on seventeen datasets covering large, small, few-shot, and AFHQ scenarios, with HP-GAN reported as consistently outperforming state-of-the-art methods on FID, plus diversity/quality measures (KID, precision, recall, PPL). Code is available at https://github.com/higun2/HP-GAN.
Main Findings
-
Ablation on FFHQ (70k) shows steady FID improvement: FastGAN 12.69 → +Projected GAN (CNN as F1) 3.39 → +ViT as F2 with small z and blur regularization 2.29 → +discriminator consistency 1.86 → +FakeTwins (HP-GAN) 1.69. Recall rises across the same progression: 0.184 → 0.464 → 0.497 → 0.524 → 0.545.
-
Ablation on Pokemon (833 images) follows the same ordering: 81.86 → 26.36 → 24.70 → 24.02 → 23.62 FID. Recall improves from 0.004 (FastGAN) to 0.259 (Config B) to 0.164 (Config C) to 0.210 (Config D) to 0.310 (HP-GAN). Perceptual path length drops from 568.5 / 574.3 (Config C) to 548.9 / 554.8 (Config D) to 450.6 / 451.3 (HP-GAN) as reported in the table.
-
Benchmark results (Table 2, FID): HP-GAN reports 1.69 on FFHQ, 1.19 on LSUN-Bedroom, and 1.44 on LSUN-Church, compared with Projected GAN (3.39 / 1.52 / 1.59), Diffusion GAN (3.73 / 1.43 / 1.85), Vision-aided GAN (3.30 / not reported / 1.72), GLeaD (2.90 / 2.72 / 2.15), ADM (2.57 / 1.90 / not reported), and LDM (4.98 / 2.95 / 4.02), among others.
-
Few-shot results (Table 3, FID): HP-GAN reports 10.30 (100-shot Obama), 13.21 (Grumpy Cat), 3.34 (100-shot Cat), 14.34 (Panda), and 15.68 (AnimalFace Dog). On 100-shot Obama, QADDRS is lower at 9.93 and DANI at 10.08; on the remaining listed few-shot columns HP-GAN's numbers are the lowest shown (e.g., Panda 14.34 vs DANI 17.72 and QADDRS 16.94; AnimalFace Dog 15.68 vs QADDRS 16.57).
-
AFHQ results (Table 4, FID): HP-GAN reports 1.81 (Cat), 3.63 (Dog), and 1.18 (Wild), compared with Projected GAN (2.16 / 4.52 / 2.17), Diffusion GAN (2.40 / 4.83 / 1.51), Vision-aided GAN (2.44 / 4.60 / 2.25), and InsGen (2.60 / 5.44 / 1.77).
-
Discriminator consistency reduces overfitting: The paper plots signed real logits sign(D(x)) — the proportion of the training set receiving positive discriminator outputs — and reports that Config D yields a lower value than Config C on FFHQ, indicating reduced discriminator overfitting; recall also improves with discriminator consistency.
-
FakeTwins loss reflects batch diversity and detail: Averaged over 100 random augmentations, the FakeTwins loss is relatively high when all images in a batch are identical (whether a simple color image or a face), lower when the batch contains varied images, intermediate when the batch holds similar images generated with noise perturbation, and lower when images have less Gaussian blur (more detail).
-
FakeTwins accelerates convergence: A plot of FID and FakeTwins loss on the Pokemon dataset shows FakeTwins (Config E) promoting faster convergence and improved FID relative to Config D.
-
CLEVR results, partially reported: In the truncated Table 5, ADA reports FID 10.17, KID 8.15, precision 0.373, recall 0.569; FastGAN reports 3.24, 2.64, 0.600, 0.650; and Projected GAN's row begins with FID 0.89 before the provided content cuts off. HP-GAN's row in this table is not shown in the available content.
Methodology in Plain English
The authors start from FastGAN and add changes one at a time so each can be measured. They first swap the single discriminator for the Projected GAN setup, which projects real and generated images into a frozen pretrained feature space (using cross-channel mixing and cross-scale mixing layers, plus four discriminators per feature pyramid). They then add a second, complementary feature network based on a Vision Transformer, following StyleGAN-XL, reduce the latent code to 64 dimensions, and blur images with a Gaussian filter of sigma = 2 pixels for the first 200k training images so the discriminators do not lock onto high-frequency detail early.
Discriminator consistency adds a squared-difference penalty between the summed CNN discriminator outputs and the summed ViT discriminator outputs on the same image, applied to both real images and generated images in the discriminator loss and to generated images in the generator loss. The intent is agreement in judgment about image quality, not feature alignment, so the two architectures keep their complementary strengths.
FakeTwins borrows Barlow Twins, which pushes the cross-correlation matrix of two embeddings toward the identity so features are decorrelated and redundant information is reduced. The twist is that the "two views" here are two augmented versions of generated images, one from G(z) and one from G(z + ε̂) with noise-based latent perturbation (ε̂ = l1 · |z|, l1 = 0.1). Those images pass through frozen pretrained projectors (CNN and ViT features, concatenated after global average pooling) into a trainable linear head with three layers of 512 units each; only the head is trained. The Barlow Twins loss (with λ1 = 0.005) is then computed on the two sets of embeddings, which means the generator itself is regularized in a rich pretrained feature space rather than only the discriminator.
Training uses hinge losses for D and G, spectral normalization without gradient penalties, exponential moving average of generator weights, the Adam optimizer, and differentiable data augmentation (also used to create FakeTwins' distorted views), with x-flips on all datasets. Hyperparameters are λ_D^f = 1, λ_D^r = 1, λ_G = 1, and λ_f = 0.02. Models train on 20M images, extended to 100M for few-shot datasets, on 8 NVIDIA RTX A5000 GPUs with batch size 64. FID is computed between 50k generated images and all training images, and the best (lowest) FID is reported.
Why This Matters
-
Impact on research: The paper argues that prior SSL-in-GAN work (ContraD, InsGen, FakeCLR, and others) treats the discriminator's feature extractor as a trainable contrastive encoder and mainly improves the discriminator, whereas HP-GAN keeps pretrained encoders frozen and applies a negative-free, information-maximization objective directly to generator outputs. It also argues that multi-backbone methods (Projected GAN, StyleGAN-XL, Vision-aided GAN) either use one backbone or treat several independently, while HP-GAN explicitly couples them with a consistency term. NoisyTwins, the closest Barlow Twins-inspired prior, is noted as label-dependent and conditional-only, whereas HP-GAN's loss operates on image features class-agnostically.
-
Real-world applications (as framed by the paper's motivation):
- Medical imaging where rare diseases yield very few training examples.
- Generating portraits of a specific set of people (the paper's "specific sets of celebrity portraits" example).
- Synthesizing artwork in the style of a particular artist (the paper's WikiArt experiment covers 1000 art paintings).
- Low-data image domains generally, including the paper's few-shot objects, animals, faces, and scene categories.
-
Industry relevance: Because HP-GAN builds on existing FastGAN generator and StyleGAN-XL discriminator components and frozen public pretrained networks (EfficientNet-lite0 and DeiT-B), it is a practical add-on to established pipelines. Its gains on small and few-shot datasets are directly relevant to settings where collecting or licensing large datasets is expensive, and its reported stability improvements target the operational pain of unreliable GAN training runs.
Future Directions
- Which pretrained backbones and how many: The paper uses EfficientNet-lite0 with DeiT-B and studies CNN-only versus CNN+ViT combinations; the optimal backbone family, number of feature networks, and whether to extend beyond a two-family CNN/ViT split remain open.
- Weighting the losses: The reported settings are λ_D^f = 1, λ_D^r = 1, λ_G = 1, and λ_f = 0.02; how sensitive results are to these choices, and whether they should adapt during training, is not established in the provided content.
- Extending beyond image generation: The paper describes FakeTwins and discriminator consistency specifically for GAN image synthesis; whether the same pretrained-encoder-as-SSL-loss idea transfers to other generative paradigms (the paper compares against diffusion-based baselines such as ADM, LDM, and Diffusion GAN) is an open question.
- Scale and resolution: Experiments resize datasets to 256² with AFHQ at 512²; whether the approach holds at higher resolutions and on the largest-scale regime (LSUN-Bedroom is 3M images) with the same hyperparameters is not reported here.
Target Audience
Researchers and graduate students working on generative adversarial networks, image synthesis, and self-supervised representation learning will get the most from this paper, particularly those interested in how frozen pretrained features can be repurposed beyond perceptual loss. It also suits practitioners facing low-data or few-shot image generation problems, and engineers who want a concrete, code-released recipe (FastGAN generator, StyleGAN-XL discriminator, hinge loss, spectral normalization, differentiable augmentation) that they can bolt onto an existing GAN pipeline. Readers without a background in GAN training objectives, Barlow Twins, or CNN/ViT architectures will find the methodology sections demanding.
Authors’ abstract
Generative Adversarial Networks (GANs) have made significant progress in enhancing the quality of image synthesis. Recent methods frequently leverage pretrained networks to calculate perceptual losses or utilize pretrained feature spaces. In this paper, we extend the capabilities of pretrained networks by incorporating innovative self-supervised learning techniques and enforcing consistency between discriminators during GAN training. Our proposed method, named HP-GAN, effectively exploits neural network priors through two primary strategies: FakeTwins and discriminator consistency. FakeTwins leverages pretrained networks as encoders to compute a self-supervised loss and applies this through the generated images to train the generator, thereby enabling the generation of more diverse and high quality images. Additionally, we introduce a consistency mechanism between discriminators that evaluate feature maps extracted from Convolutional Neural Network (CNN) and Vision Transformer (ViT) feature networks. Discriminator consistency promotes coherent learning among discriminators and enhances training robustness by aligning their assessments of image quality. Our extensive evaluation across seventeen datasets-including scenarios with large, small, and limited data, and covering a variety of image domains-demonstrates that HP-GAN consistently outperforms current state-of-the-art methods in terms of Fréchet Inception Distance (FID), achieving significant improvements in image diversity and quality. Code is available at: https://github.com/higun2/HP-GAN.