Research
Composite Classifier-Free Guidance for Multi-Modal Conditioning in Wind Dynamics Super-Resolution
Overview Research area: Machine learning for weather and climate — specifically generative diffusion models applied to wind dynamics super-resolution, combining numerical weather prediction (NWP) data

- arXiv
- 2512.13729
- Published
- 2025-12-13
- Authors
- Jacob Schnell, Aditya Makkar, Gunadi Gani, Aniket Srinivasan Ashok, Darren Lo, Mike Optis, Alexander Wong, Yuhao Chen
AI summary
Overview
Research area: Machine learning for weather and climate — specifically generative diffusion models applied to wind dynamics super-resolution, combining numerical weather prediction (NWP) data with deep learning.
Technical level: Advanced. The paper assumes familiarity with diffusion probabilistic models, score functions, and classifier-free guidance, though the core idea can be understood at a conceptual level.
Scope: This paper introduces a composite generalization of classifier-free guidance (CCFG) for diffusion models with many conditioning inputs, and applies it to WindDM, a 100M-parameter U-Net diffusion model that super-resolves coarse ERA5 wind data to 3 km WRF-resolution wind fields.
What This Paper Is About
Producing high-resolution, accurate wind data is expensive: classical dynamical NWP models like WRF are the gold standard but cost thousands of dollars per run, while cheap statistical methods are much less accurate. Deep learning has been proposed as a middle ground, but wind super-resolution differs from natural image super-resolution — it uses far more than the usual 3 RGB input channels (WRF uses 293 variables, the GAN-based Sup3r model uses 16, and WindDM's "all" configuration uses 8 input variables), and existing conditional diffusion work mostly targets the single-conditioning-variable case. The paper asks how to better exploit multiple conditioning modalities in diffusion models, and answers with a new inference scheme plus a purpose-built model.
Key Contributions
-
Composite classifier-free guidance (CCFG): A generalization of standard CFG to multiple conditioning inputs. Instead of up-weighting one complete-likelihood term, CCFG samples from a composite likelihood built from subsets of the conditioning variables, each with its own weight. The standard CFG case is recovered when the number of subsets m = 1 and the single subset equals the full conditioning set.
-
A drop-in property: CCFG can be applied to any pre-trained diffusion model that was trained with standard CFG dropout. WindDM was trained with each conditioning variable independently dropped out with probability p = 0.1, which is the straightforward generalization of usual CFG dropout to multiple variables.
-
A CCFG model selection algorithm (Algorithm 1): A gradient-descent procedure that jointly selects which conditioning subsets K₁ … K_m to use and their weights w, constrained to a simplex summing to a user-specified total weight W, with L₁ and L₂ weight decay on w, periodic greedy pruning of the least-weighted subset, and projection back to the W-simplex. This introduces m as a compute budget parameter trading model evaluations for sample quality.
-
WindDM, a diffusion model for industrial-scale wind reconstruction: Trained on a new dataset of 265,390 timestamps of paired low-resolution and high-resolution wind data with 8 input variables and 2 target variables, achieving state-of-the-art reconstruction quality among deep learning models at up to 1000× lower cost than classical methods.
Main Findings
-
WindDM leads on the primary benchmark: On the United Kingdom domain (trained on the other 6 domains, evaluated on 2022 data), WindDM (all variables) achieves a mean-map RMSE of 0.437 and WindDM (basic) achieves a timestamp RMSE of 2.142. For context, the U-Net baseline scores 0.452 mean-map RMSE and 2.257 timestamp RMSE, Sup3r scores 0.707 and 2.569, CorrDiff (all) scores 1.135 and 3.942, Random Forest scores 0.592 and 2.749, and bicubic interpolation scores 1.082 and 2.246. On CRPS, WindDM (basic) reaches 0.415 mean-map and 1.291 timestamp, versus 0.347 and 1.750 for U-Net and 2.855 for CorrDiff (basic) on timestamps.
-
CCFG improves on CFG, and both beat direct inference: For the basic variable setup, direct inference gives 0.608 mean-map RMSE / 2.909 timestamp RMSE, CFG gives 0.493 / 2.413, and CCFG gives 0.534 / 2.142. For the all-variables setup, direct gives 0.305 / 2.643, CFG gives 0.523 / 2.433, and CCFG gives 0.437 / 2.348. The paper describes CCFG's improvement over CFG as modest but consistent in timestamp RMSE.
-
CCFG is complementary to ensembling: Excluding ensembling, direct inference uses 1 neural function evaluation (NFE) per reverse-process step, CFG uses 2, and CCFG with m = 2 uses 4. The authors find CCFG's performance does not saturate as rapidly as CFG or direct inference when ensemble size is increased.
-
Cheap relative to dynamical models: Even at 16 NFEs per inference step, evaluating the entire United Kingdom domain takes 4 GPU hours on an NVIDIA L40S, costing an effective $7.44 USD under AWS on-demand EC2 pricing as of September 19th, 2025, versus several thousand dollars for a comparable WRF run.
-
Generalization across domains: In a 7-fold cross-validation over UK, Italy, Spain, Switzerland, Northern Sweden, Southern Sweden, and the Norwegian Sea, WindDM consistently shows the best generalization of the compared models, including to the offshore Norwegian Sea domain despite training only on onshore data.
-
The selection algorithm matters: Ablating Algorithm 1 on the basic model gives 0.534 mean-map RMSE / 2.142 timestamp RMSE at 4 NFEs. Removing weight decay gives 0.503 / 2.203 (4 NFEs), removing pruning 0.497 / 2.227 (8 NFEs), replacing simplex projection with normalization 0.499 / 2.213 (4 NFEs), uniform weights over subsets 0.481 / 2.382 (8 NFEs), the exhaustively searched best subsets 0.566 / 3.091 (5 NFEs), and normal CFG 0.483 / 2.413 (2 NFEs).
-
Channel concatenation beats other conditioning methods: With normal CFG, U-Net with concatenation scores 0.483 / 2.413; U-Net with cross-attention 0.641 / 2.810; DiT with concatenation 0.504 / 2.301; DiT with cross-attention 0.772 / 2.906; and DiT with AdaNorm 0.704 / 2.348. The DiT/concatenate model took 3× as long to train as U-Net/concatenate, which is why U-Nets were used elsewhere.
Methodology in Plain English
The authors start from SR3, a diffusion super-resolution design that denoises a noisy high-resolution image while the low-resolution image is concatenated as extra input channels. WindDM extends this by concatenating additional guiding maps (such as topography and temperature) alongside the low-resolution wind data, and by adding three auxiliary losses alongside the main L₁ denoising objective — a physics-informed neural network (PINN) loss, a wavelet-domain loss, and a Sobel filter loss — aimed at fine-grain detail in complex terrain. The authors note they also tried predicting the noise directly, which performed worse empirically.
For guidance, standard CFG works by comparing a fully conditioned prediction against a fully unconditioned one, effectively up-weighting samples the model thinks are likely under the conditioning. The authors argue this implicit likelihood estimate is unlikely to match the true likelihood when conditioning data are highly complex, as in wind super-resolution. Their alternative is to decompose the conditioning set into subsets, run a partially conditioned prediction for each subset, and sum the weighted differences between those predictions and the unconditioned prediction. Because composite likelihoods often have smoother surfaces than complete likelihoods while remaining unbiased, the authors expect more stable guidance. The number of subsets m acts as a compute dial.
To pick which subsets and weights to use, they restrict the search to subsets missing at most p conditioning variables, giving a search space of size n_p ∈ O(k^p) rather than 2^k − 1 non-empty subsets. Starting from uniform weights summing to W, they run gradient descent on an L₁ reconstruction loss plus L₁ and L₂ penalties on the weights, periodically prune the lowest-weighted subset, and reproject onto the W-simplex. Because the diffusion model itself is frozen and only the weights are learned, only the weight vector is optimized.
Inference uses the DPM++ sampler with a 3rd order multistep ODE solver and 10 inference steps, on a 4-layer convolutional U-Net with 2 self-attention layers and 10% dropout, trained with quantized 16-bit weights for 50 epochs (72 GPU minutes per epoch on a single NVIDIA L40S for the all-variables model), using a DDPM forward schedule with T = 1000 steps and betas linearly increasing from 0.0001 to 0.02. The model selection step uses a cheaper 5-step DDPM sampler and gradient checkpointing.
Why This Matters
Impact on research: The paper reframes multi-variable conditioning in diffusion models as a composite-likelihood problem rather than a single implicit-classifier problem. Because CCFG drops into any pre-trained model trained with CFG dropout and introduces an explicit compute-for-quality knob, it is a general-purpose idea that could apply wherever models condition on many correlated inputs. The paper also documents concrete ways in which wind super-resolution departs from natural image super-resolution: the low-resolution and high-resolution distributions come from different physics models rather than from coarsening, downstream goals are distribution-level (annual mean maps) rather than per-pixel, and input channel counts are far higher.
Real-world applications:
- Wind farm siting and turbine micro-siting: Small placement errors can cost several megawatts of missed energy capture over time, so high-resolution wind maps inform where turbines go.
- Annual energy production and resource assessment: The annual mean map, produced by averaging a full year of predictions, is described as a common downstream product for wind farm planning.
- Cheap replacement or augmentation of dynamical downscaling: Reconstructing mesoscale (1–3 km) fields from global-scale data (0.25° ≈ 30 km) for a fraction of the cost of a WRF run.
- Offshore and complex-terrain assessment: The model generalizes to the offshore Norwegian Sea domain and to coastal and mountainous UK terrain despite training on other domains.
Industry relevance: Renewables supplied 25.5% of global renewable energy capacity from wind turbines per the paper's citation, and the cost asymmetry between neural methods ($7.44 effective per full-domain evaluation versus several thousand dollars for WRF) is directly relevant to wind developers and consultancies. Two authors are affiliated with Veer Renewables, indicating industrial interest in the tooling.
Future Directions
- Pushing the NFE budget further: Since CCFG performance did not saturate as rapidly as CFG or direct inference under increasing ensemble size, it remains open how much quality is available at higher compute budgets, and where that curve flattens.
- Conditioning architecture trade-offs: DiT models marginally improved performance over U-Net on this task but took 3× longer to train. Whether the quality gain justifies the cost at larger scales, and whether concatenation remains optimal at higher input counts, is unresolved.
- Applying CCFG outside wind: The paper motivates CCFG specifically by wind's many conditioning variables, but the method is domain-agnostic. The related work discusses non-natural image domains such as remote sensing and medical imaging, which are natural candidates.
- Hyperparameter selection for CCFG: The paper fixes p (exclusion count), W = 1.5, and m per configuration, and the search space is restricted to subsets missing at most p variables. How to choose these in general, and whether richer subset structures help, is left open.
Target Audience
This paper will be most useful to machine learning researchers working on diffusion models with multi-conditioning inputs, and to applied scientists and engineers building data-driven downscaling or super-resolution systems for weather and wind resource assessment. Practitioners evaluating whether to replace or supplement expensive dynamical NWP runs with neural surrogates will also find the cost and accuracy comparisons directly relevant. Readers need a working understanding of diffusion sampling and guidance to follow the method section, though the results tables are legible without it.
Authors’ abstract
Various weather modelling problems (e.g., weather forecasting, optimizing turbine placements, etc.) require ample access to high-resolution, highly accurate wind data. Acquiring such high-resolution wind data, however, remains a challenging and expensive endeavour. Traditional reconstruction approaches are typically either cost-effective or accurate, but not both. Deep learning methods, including diffusion models, have been proposed to resolve this trade-off by leveraging advances in natural image super-resolution. Wind data, however, is distinct from natural images, and wind super-resolvers often use upwards of 10 input channels, significantly more than the usual 3-channel RGB inputs in natural images. To better leverage a large number of conditioning variables in diffusion models, we present a generalization of classifier-free guidance (CFG) to multiple conditioning inputs. Our novel composite classifier-free guidance (CCFG) can be dropped into any pre-trained diffusion model trained with standard CFG dropout. We demonstrate that CCFG outputs are higher-fidelity than those from CFG on wind super-resolution tasks. We present WindDM, a diffusion model trained for industrial-scale wind dynamics reconstruction and leveraging CCFG. WindDM achieves state-of-the-art reconstruction quality among deep learning models and costs up to $1000\times$ less than classical methods.