Skip to content
AI.info

Research

VeCoR -- Velocity Contrastive Regularization for Flow Matching

VeCoR — Velocity Contrastive Regularization for Flow Matching Overview Research area: Generative modeling / computer vision — specifically Flow Matching (FM) and its training objectives for image synt

VeCoR -- Velocity Contrastive Regularization for Flow Matching
arXiv
2511.18942
Published
2025-11-24
Authors
Zong-Wei Hong, Jing-lun Li, Lin-Ze Li, Shen Zhang, Yao Tang

AI summary

VeCoR — Velocity Contrastive Regularization for Flow Matching

Overview

Research area: Generative modeling / computer vision — specifically Flow Matching (FM) and its training objectives for image synthesis.

Technical level: Intermediate to Advanced. The paper assumes familiarity with diffusion models, probability-flow ODEs, stochastic interpolants, and latent-space generative training.

Scope: A one-paragraph-style summary of a training-scheme paper that adds a contrastive "negative velocity" term to the standard flow-matching loss and evaluates it on ImageNet-1K class-conditional generation and MS-COCO text-to-image generation.

What This Paper Is About

Standard Flow Matching trains a model to predict a velocity field that moves samples from noise toward data, but it only tells the model where to go — it never says where not to go. The authors argue that under lightweight models or low-step sampling, small errors in the learned velocity field accumulate along the trajectory and push samples slightly off the data manifold, producing desaturated colors, geometric distortions, blur, and hallucinated artifacts. The goal is to add a complementary repulsive training signal that keeps trajectories anchored to the data manifold, without adding any new network, data, or architectural change.

Key Contributions

  1. A complementary training scheme for flow-based models that augments standard supervision with an ensemble of stable and perturbed flows, improving sample quality and convergence without extra data or architectural changes.
  2. Velocity Contrastive Regularization (VeCoR), a contrastive loss defined directly on the velocity field that enforces directional consistency of generative trajectories, yielding more stable and faster training.
  3. A general negative-candidate construction mechanism using augmentation-like perturbations applied across three representational domains — image, latent, and velocity — that preserve semantic consistency while producing "plausible yet dynamically inconsistent" negative velocities.
  4. Empirical results across scales and benchmarks: on ImageNet-1K 256×256, 22% and 35% relative FID reductions on SiT-XL/2 and REPA-SiT-XL/2 backbones respectively, and a further 32% relative FID reduction on MS-COCO text-to-image generation, with consistent gains especially in low-step and lightweight settings.

Main Findings

  • Class-conditional gains scale with model size: On ImageNet-1K 256×256 (50 NFEs, Euler–Maruyama), VeCoR reduced SiT-S/2 FID from 64.26 to 55.13, SiT-B/2 from 42.28 to 33.30, SiT-L/2 from 24.06 to 18.86, and SiT-XL/2 from 20.01 to 15.56. The paper reports FID reductions of 14–22% and sFID reductions of 44–53% across these backbones, with recall largely preserved.
  • Comparison against Delta FM: At SiT-B/2, VeCoR (FID 33.30) is roughly on par with ΔFM (33.39); at SiT-XL/2, VeCoR is stronger (FID 15.56 vs. 16.32, with higher IS of 80.96 vs. 78.07).
  • Generalization to REPA backbones: With REPA-SiT, VeCoR reduced FID from 27.33 to 20.39 at B/2 and from 11.14 to 7.28 at XL/2 (a 25–35% relative reduction), improved sFID from 11.70 to 5.57 and from 8.25 to 5.17 (37–52% improvement), and raised IS from 61.60 to 69.09 and from 115.83 to 127.90.
  • Text-to-image on MS-COCO: Using the MMDiT+REPA pipeline at 150K iterations, batch size 256, hidden dimension 768, and depth 24, VeCoR with Random Crop and Resize reached FID 4.82 with the ODE (Heun) solver and 4.55 with the SDE (Euler–Maruyama) solver at CFG scale 2.0, versus 5.16 and 4.82 for the reproduced ΔFM baseline. Under lower guidance (SDE, ω=1.0), VeCoR with Random Channel Shuffle cut FID from 9.87 to 6.65.
  • Combining with classifier-free guidance: After a grid search over w ∈ {1.25, 1.75, 1.8, 1.85, 2.25}, σ_low = 0, and σ_high ∈ {0.50, 0.65, 0.75, 1.0}, VeCoR achieved FID 1.94 and sFID 4.45 at identical optimal hyperparameters (w=1.85, σ_low=0.0, σ_high=0.65), slightly surpassing ΔFM (FID 1.97, sFID 4.49) and REPA SiT-XL/2 (FID 2.09, sFID 5.55).
  • Perturbation domain matters: Ablations with SiT-S/2 (baseline FID 64.26) show that velocity-space perturbations generally beat image- and latent-space ones. Random Channel Shuffle in velocity space gave FID 55.13; Random Crop & Resize gave 59.43; CutMix gave 59.62; Gaussian Blur gave 58.73; Gaussian Noise gave 65.07; Color Jitter gave 64.69.
  • Negative-set size: Channel Shuffle with K=2 gave the best trade-off (FID 52.60, IS 26.21, sFID 6.89, Precision 0.420, Recall 0.596). K=1 gave FID 55.13; K=3 gave 53.96; K=4 gave 53.70. Random Crop was consistently weaker (59.43 at K=1, 59.34 at K=2) and K-insensitive, and Shuffle+Crop combined gave FID 56.16, worse than Channel Shuffle alone.
  • The regularization weight is a trade-off: λ too small (e.g., 0.01) leaves noticeable artifacts; λ ≈ 0.05 gives the best FID and sFID with sharper images; λ in the 0.1–0.2 range over-constrains the model and loses fine-grained detail.
  • Default configuration: negatives formed in velocity space via random channel shuffling with K=1 and fixed contrastive weight λ=0.05.
  • Faster convergence and better low-step sampling: SiT-XL/2+VeCoR converges faster to lower FID over training epochs, and achieves better FID at low NFE (≤ 50) while remaining comparable at higher NFE, without asymptotic degradation.

Methodology in Plain English

The starting point. Flow Matching defines a path that interpolates between noise and data using time-dependent schedules α_t and σ_t (the paper uses the linear setting α_t = t and σ_t = 1 − t). The model learns a neural network v_θ(x̂_t, t) to match a ground-truth target velocity v̂ = α̇_t x̂ + σ̇_t ε, minimizing mean-squared error over a finite training set of (x̂, t, ε) triples. Images are first encoded into latents (32×32×4) with the Stable Diffusion VAE.

The added ingredient. The authors treat the velocity itself as editable data. For each training sample they produce one positive velocity (the ground-truth target) and K negative velocities that are semantically plausible but dynamically wrong. The loss then both attracts the prediction toward the positive and subtracts a weighted term pushing it away from each negative, with weight λ ∈ (0, 1):

L = (1/N) Σ [ ||v_θ − v̂₊||² − λ Σⱼ ||v_θ − v̂₋⁽ⁱʲ⁾||² ]

Where negatives come from. Rather than mining hard negatives from real data (costly and ill-defined), the authors borrow the augmentation taxonomy of SimCLR: spatial/geometric transformations (random cropping and resizing, channel shuffling, CutMix) and appearance transformations (color jitter, Gaussian blur, additive Gaussian noise). They apply these not only in image space but reinterpret them for latent and velocity tensors. For example, random channel shuffle applies a per-sample cyclic shift k ∈ {1, …, C−1} so no channel stays in place; random crop and resize samples an area ratio α ∈ [0.9, 0.95] and aspect ratio r ∈ [0.95, 1.05], crops, and resizes back; CutMix uses a derangement permutation so no sample mixes with itself. Whichever domain the negative comes from, the authors treat it uniformly as v̂₋.

Why it should work. The contrastive term is described as steering predicted velocities away from the mean of synthesized off-manifold trajectories — meaning the model learns to suppress drift rather than only pursue a target direction. Because this conflicts with naive classifier-free guidance, which steers away from the unconditional flow, the authors adopt the ΔFM strategy of folding the pre-computed mean of the off-manifold trajectories into the guidance equation.

Evaluation. Class-conditional experiments use ImageNet-1K at 256×256 following ADM preprocessing, training SiT backbones at S/2, B/2, L/2, and XL/2 scales under identical hyperparameters except for VeCoR, and separately REPA-SiT at B/2 and XL/2 (REPA aligns intermediate representations with a pretrained DINOv2 encoder). Metrics are FID, IS, sFID, Precision, and Recall on 50,000 generated samples, using the SDE Euler–Maruyama sampler with w_t = σ_t and 50 NFEs. Text-to-image experiments follow the U-ViT setup, training REPA-MMDiT from scratch on MS-COCO with CLIP text embeddings.

Why This Matters

Impact on research. VeCoR reframes flow-matching supervision from a one-sided attraction problem into a two-sided attract–repel problem, drawing an explicit contrast with the concurrent ΔFM work: ΔFM targets inter-class semantic discriminability by pushing away from the data-averaged expectation, whereas VeCoR targets geometric stability within the vector field by repelling corruption-induced drift. It is plug-and-play — no new networks, no external data — which makes it easy to test alongside existing FM, rectified-flow, and distillation pipelines.

Real-world applications (implied by the paper's emphasis on low-step and lightweight settings):

  • Fast on-device or edge image generation where only a small number of sampling steps is affordable.
  • Latent-space text-to-image systems such as the MMDiT+REPA pipeline evaluated on MS-COCO.
  • Small-model deployments, since the paper reports the largest relative FID gains on smaller backbones (SiT-S/2 and B/2).
  • Any generative pipeline where perceptual artifacts — color shifts, geometric distortion, blur, or hallucinated structures — are the main failure mode rather than gross semantic errors.

Industry relevance. The method's selling points are exactly the ones that matter for production: training stability, faster convergence under the same computational budget, and better quality at low NFE. The authors' affiliations (JIIOV Technology) and their use of widely adopted backbones (SiT, REPA, Stable Diffusion VAE, CLIP) suggest the target is practical deployment rather than a bespoke architecture.

Future Directions

  1. Adaptive hard-negative mining. The authors explicitly state that the current negative sampling strategy remains heuristic and data-agnostic, and name adaptive hard-negative mining as future work.
  2. Trajectory-aware perturbations. Also named by the authors as a direction to strengthen stability and efficiency beyond the current augmentation-like perturbations.
  3. Beyond augmentation-based negatives. The paper notes the framework is not limited to augmentation and could support domain adaptation, consistency regularization, or adversarial robustness.
  4. Open questions the paper leaves unaddressed: how to select λ and K without a grid search, whether appearance-based and geometric perturbations could be combined more effectively than the failed Shuffle+Crop pairing, and whether the gains persist for larger-scale or non-image modalities (the paper reports no such experiments).

Target Audience

Researchers and engineers working on flow matching, diffusion models, or continuous normalizing flows who want a training-time improvement they can drop into an existing codebase. It is most useful to practitioners focused on low-step sampling, small or lightweight backbones, and latent-space image or text-to-image generation. Readers without prior exposure to stochastic interpolants, probability-flow ODEs, and standard diffusion evaluation metrics will find Sections 3 and 5.1 difficult without supplementary reading.

Authors’ abstract

Flow Matching (FM) has recently emerged as a principled and efficient alternative to diffusion models. Standard FM encourages the learned velocity field to follow a target direction; however, it may accumulate errors along the trajectory and drive samples off the data manifold, leading to perceptual degradation, especially in lightweight or low-step configurations. To enhance stability and generalization, we extend FM into a balanced attract-repel scheme that provides explicit guidance on both "where to go" and "where not to go." To be formal, we propose \textbf{Velocity Contrastive Regularization (VeCoR)}, a complementary training scheme for flow-based generative modeling that augments the standard FM objective with contrastive, two-sided supervision. VeCoR not only aligns the predicted velocity with a stable reference direction (positive supervision) but also pushes it away from inconsistent, off-manifold directions (negative supervision). This contrastive formulation transforms FM from a purely attractive, one-sided objective into a two-sided training signal, regularizing trajectory evolution and improving perceptual fidelity across datasets and backbones. On ImageNet-1K 256$\times$256, VeCoR yields 22\% and 35\% relative FID reductions on SiT-XL/2 and REPA-SiT-XL/2 backbones, respectively, and achieves further FID gains (32\% relative) on MS-COCO text-to-image generation, demonstrating consistent improvements in stability, convergence, and image quality, particularly in low-step and lightweight settings. Project page: https://p458732.github.io/VeCoR_Project_Page/

Read the original paper