Research
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
Overview Research area: Computer vision and generative modeling — specifically inference-time guidance for diffusion and flow-matching generative models (diffusion transformers such as DiT-XL/2 and MM

- arXiv
- 2512.17303
- Published
- 2025-12-19
- Authors
- Ankit Yadav, Ta Duc Huy, Lingqiao Liu
AI summary
Overview
Research area: Computer vision and generative modeling — specifically inference-time guidance for diffusion and flow-matching generative models (diffusion transformers such as DiT-XL/2 and MMDiT / Stable Diffusion 3 Medium).
Technical level: Intermediate. The paper assumes familiarity with diffusion and flow-matching sampling, classifier-free guidance (CFG), and self-attention in transformers, but the core idea is explained in accessible terms.
One-sentence scope: The paper proposes EMAG, a training-free method that replaces selected self-attention maps with their exponential moving average at inference time to create controllable, fine-grained "hard negatives" that steer sampling toward higher-quality outputs.
What This Paper Is About
Modern diffusion and flow-matching models rely on guidance techniques — most commonly classifier-free guidance (CFG) — to improve sample quality and prompt adherence. A newer family of methods steers generation by contrasting a strong model with a deliberately weakened one, but existing approaches offer limited control over how degraded that weaker signal is, and they typically target fixed layers. This paper introduces Exponential Moving Average Guidance (EMAG), which regulates the difficulty of negative samples through a single decay factor and selects which attention layers to perturb at each timestep using a statistics-based rule.
Key Contributions
-
EMAG, a training-free attention-EMA guidance mechanism: The method perturbs self-attention by replacing the attention map at a selected layer and timestep with its exponential moving average, producing controllable weak signals without any additional training. The EMA decay factor β governs how much high-frequency detail the weaker branch destroys, and therefore how hard the negative sample is.
-
A quantitative study of negative-sample hardness: The authors measure hardness as cosine similarity between the negative sample and the final guided output in frozen CLIP and DINOv2 embedding spaces, showing that EMAG produces harder negatives than prior guidance methods, that SAG is the closest competitor, and that smaller β within EMAG yields harder negatives.
-
Statistics-based adaptive layer selection: Rather than fixing the target layers, EMAG computes the mean absolute error (MAE) between the EMA and the current attention map for each candidate layer, and at each timestep replaces the layer with the largest discrepancy. For DiT-XL/2 the candidate range is layers (12, 15); for SD3-Medium it is layers (6, 8).
-
Extensive evaluation and composition experiments: Experiments cover class-conditional and text-to-image generation on DiT-XL/2 (256×256 and 512×512) and MMDiT via Stable Diffusion 3 Medium (1024×1024), reporting competitive results against CFG, SAG, PAG, SEG, ERG, S²-Guidance, APG, and CADS, plus evidence that EMAG composes with APG and CADS.
Main Findings
-
HPS gain over CFG: EMAG improves the Human Preference Score (HPS v2, scaled by 100) by +0.54 over CFG, reaching 29.76 ± 0.03 versus CFG's 29.22 ± 0.05 on the SD3 COCO 2014 text-to-image setting. Its FID was 22.890 versus CFG's 22.877.
-
EMAG as a plug-in to other guidance methods: In the SD3 COCO 2014 comparison, adding EMAG raised CFG from 29.22 to 29.76, ERG from 29.35 to 29.56, S² from 29.15 to 29.51, and SEG + CFG from 29.25 to 29.29. SAG + CFG is listed at 29.28 with no EMAG combination reported.
-
Combination with orthogonal guidance: EMAG + APG achieved the highest reported HPS at 29.79 ± 0.04 (FID 21.819), and EMAG-Q + APG reached 29.77 (FID 22.480). EMAG + CADS reached 28.86 ± 0.02 (FID 19.950), and EMAG-Q + CADS 28.96 (FID 20.030).
-
Variants of EMAG: EMAG-I (perturbing only the Image→Image sub-block in MMDiT joint attention) scored HPS 29.56 with FID 22.154. EMAG-Q (tracking EMA over query embeddings instead of post-softmax attention maps) scored HPS 29.64 with FID 23.530. PAG + CFG is marked as incompatible with SD3.
-
Unconditional generation gains: On SD3 (MMDiT) unconditional COCO 2014, EMAG reached the lowest FID at 74.98 (versus 101.94 for no guidance, 95.38 for SAG, 90.21 for SEG, 82.11 for I-ERG, and 111.24 for the PAG-incompatible entry) and the highest Coverage at 0.112. On DiT-XL/2 256×256, EMAG reached FID 29.85 and Coverage 0.768 (versus 52.76 and 0.554 for no guidance), with Precision 0.621 and Recall 0.561.
-
Layer selection matters: In a 5K-sample SD3 ablation, fixed single layers gave FID/HPS of 27.47/29.54 (L6), 28.94/29.51 (L7), and 28.23/29.66 (L8), perturbing all layers gave 35.60/28.80, and adaptive selection gave 28.52/29.60 — described as the best FID–HPS balance.
-
Speed–quality trade-off: On a single NVIDIA A100-SXM4-40GB with SD3 Medium at 1024×1024, 28 steps, batch size 1 (averaged over 45 images after 5 warmup iterations), CFG took 4.04 s/image at 16.89 GB peak memory. EMAG took 12.66 s/image (+213.2% overhead, 29.58 GB), while EMAG-Q took 7.86 s/image (+94.5%, 16.96 GB). EMAG-Q therefore cuts overhead from +213% to +95% and peak memory from 29.58 GB to 16.96 GB, at a cost of +0.63 FID and −0.12 HPS relative to EMAG (Table S16).
-
Hardness does not determine quality alone: The paper notes that S² and ERG show comparable negative hardness yet differ on the Pareto frontier, indicating that the perturbation mechanism itself also matters.
-
Default hyperparameters: β = 0.988, set via EMA half-life H through β = e^(−ln(2)/H); guidance becomes net-positive once the half-life exceeds approximately 0.7T of the sampling horizon, and the default corresponds to a half-life of approximately 2T for 28 steps. λ = 1 (hard replace) is the default. Unless noted, a CFG scale of 7 and an EMAG scale of 1.5 are used for conditional generation and 5.125 for unconditional. EMAG is applied in a late-step "tail" window with τ_s = t_max and τ_e = ⌊0.2·t_max⌋, giving τ_e = 5 for SD3 with 28 steps and τ_e = 50 for DiT with DDIM-250.
-
Evaluation protocol: 50K samples for class-conditional ImageNet experiments and 40K samples for text-to-image (one per caption from the COCO-2014 validation split), seed 8. DiT-XL/2 uses DDIM with 250 deterministic steps (η = 0); SD3 uses the FlowMatchEulerDiscrete sampler with 28 steps. Metrics are FID, HPS v2, Precision/Density, and Recall/Coverage, with CLIPScore reported for text–image alignment.
Methodology in Plain English
Diffusion models generate images by repeatedly denoising random noise, and guidance techniques steer this process by comparing two predictions — usually a conditional one and an unconditional one. EMAG instead builds a "weaker" version of the same model by tampering with its self-attention during sampling: at the chosen layer and timestep, the current attention map is swapped for a running exponential average of attention maps seen so far.
Because attention refines images from coarse to fine, this swap suppresses the fine-grained, high-frequency refinements while leaving the overall structure of the image intact. The result is a subtly degraded prediction — a hard negative — that still looks semantically plausible. The authors then contrast this degraded prediction with the original one (a first guidance update), and feed the result into a standard CFG-style update. The decay factor β controls how slowly the average moves, which in turn controls how much detail gets destroyed; smaller β means a faster-moving average and a harder negative.
To decide where to intervene, the method maintains EMA buffers for a candidate band of middle layers and, at each timestep, picks the layer whose EMA differs most from its current attention map, measured by mean absolute error. Only that one layer is replaced (λ = 1). The authors also describe two efficiency variants: EMAG-I, which limits perturbation to the image-to-image block, and EMAG-Q, which tracks the EMA over query embeddings instead, allowing fused scaled dot-product attention and reducing overhead.
For evaluation, the authors sweep hyperparameters for every baseline to construct FID–HPS Pareto frontiers and select configurations from those frontiers, then compare on identical data, sampling steps, and evaluation code.
Why This Matters
Impact on research: The paper argues that the type and difficulty of negative samples matters, not just their existence, and provides a measurable way to quantify hardness via embedding cosine similarity. It also shows that attention-EMA perturbation is complementary to orthogonal guidance strategies (APG, CADS) and to negative-signal methods (S²), suggesting a composable building block rather than a replacement for CFG.
Real-world applications:
- Text-to-image generation products where human preference scores and fine-detail fidelity (faces, hands, small objects, text in images) drive user satisfaction.
- Class-conditional image synthesis pipelines using DiT-style backbones, where post-training guidance improvements require no retraining.
- Creative and design tools that need higher perceptual quality without retraining or fine-tuning large models.
- Content generation systems where a quality boost must be delivered under limited inference budgets — EMAG-Q is presented as the lower-overhead option.
Industry relevance: Because EMAG is training-free and model-agnostic in principle, it can be dropped into existing diffusion-transformer deployments without new training data or weights. The paper is candid about the cost: the default EMAG carries +213.2% inference overhead and 29.58 GB peak memory at 1024×1024, while EMAG-Q reduces this to +94.5% and 16.96 GB. Whether those costs are acceptable is a deployment decision the paper frames directly.
Future Directions
- Moving beyond greedy layer selection: The authors state that the statistic-guided selection could be replaced by a carefully designed optimization over the per-layer discrepancy, rather than the greedy argmax used here, and defer this to future work.
- Understanding what drives quality beyond hardness: The observation that S² and ERG have comparable hardness but different Pareto behavior implies that the perturbation mechanism itself needs further study.
- Reducing the overhead of the full EMAG formulation: EMAG-Q narrows the gap but still costs +94.5% over CFG and trails EMAG slightly in HPS, leaving room for methods that preserve the post-softmax EMA's benefits at lower cost.
- Extending controllability: The paper notes that β controls the granularity of degradations, which raises the open question of how far this controllability can be pushed for more targeted refinement.
Target Audience
Researchers and practitioners working on diffusion and flow-matching generative models — particularly those studying inference-time guidance, diffusion transformers, and text-to-image quality. It is most useful to readers already comfortable with CFG and self-attention mechanics who want a training-free, drop-in quality improvement with an explicit knob for controlling negative-sample difficulty, and who are willing to weigh the reported compute overhead against the reported HPS gains.
Authors’ abstract
In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free guidance (CFG) is the de facto choice in modern systems and achieves this by contrasting conditional and unconditional samples. Recent work explores contrasting negative samples at inference using a weaker model, via strong/weak model pairs, attention-based masking, stochastic block dropping, or perturbations to the self-attention energy landscape. While these strategies refine the generation quality, they still lack a reliable control over the granularity or difficulty of the negative samples, and target-layer selection is often fixed. We propose Exponential Moving Average Guidance (EMAG), a training-free mechanism that modifies attention at inference time in diffusion transformers, with a statistics-based, adaptive layer-selection rule. Unlike prior methods, EMAG produces harder, semantically faithful negatives (fine-grained degradations), surfacing difficult failure modes, enabling the denoiser to refine subtle artifacts, boosting the quality and human preference score (HPS) by +0.54 over CFG. We further demonstrate that EMAG naturally composes with advanced orthogonal guidance techniques, such as APG and CADS, further improving HPS.