Research
ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation
Overview Research area: Computer vision, specifically cross-modal remote sensing image translation (SAR-to-EO), building on latent generative modeling, flow matching, and vision foundation model repre

- arXiv
- 2609.00968
- Published
- 2026-09-01
- Authors
- Jeonghyeok Do, Seungchul Lee, Munchurl Kim
AI summary
Overview
Research area: Computer vision, specifically cross-modal remote sensing image translation (SAR-to-EO), building on latent generative modeling, flow matching, and vision foundation model representation alignment.
Technical level: Advanced. The paper assumes familiarity with variational autoencoders, diffusion and flow-matching objectives, Diffusion Transformers (DiT), classifier-free guidance, and representation alignment losses.
Scope: The paper introduces ReFlowSET, a conditional latent flow-matching framework for SAR-to-EO translation that selects its autoencoder through a joint SAR–EO reconstruction audit and trains a conditional DiT from scratch with training-only EO representation alignment.
What This Paper Is About
SAR imagery can be captured regardless of illumination and under most weather conditions, but its speckle and geometric distortions make it hard to interpret compared with electro-optical (EO) imagery. SAR-to-EO image translation (SET) tries to generate an interpretable EO image from a corresponding SAR observation, which is an inherently ambiguous mapping because SAR backscatter and EO reflectance describe different physical properties. The paper argues that existing latent diffusion methods silently inherit an autoencoder pretrained on natural images, and that this codec choice — which controls how much SAR structure survives for conditioning and how faithfully the EO target can be reconstructed — should be an explicit design decision rather than a fixed implementation detail.
Key Contributions
-
Systematic modality-wise codec audit. The authors evaluate the pretrained autoencoders of SD2.1, SDXL, SD3.0/3.5, and FLUX.1/2 on QXS-SAROPT, SAR2Opt, SpaceNet6, and SAR-1M, measuring encode–decode round-trip PSNR for SAR and EO separately on training splits, and identify FLUX.2 as the least distortion-limited latent space. They state this is the first systematic modality-wise analysis of pretrained autoencoder upper-bounds for SET.
-
A latent flow-matching framework with dual-stream SAR conditioning. ReFlowSET trains a conditional DiT from scratch using conditional flow matching, with separately parameterized SAR and noisy-EO streams for the first blocks that are later fused channel-wise and refined jointly, instead of the input-level concatenation used in prior latent SET models.
-
Training-only EO representation alignment. Intermediate noisy-EO DiT features are aligned by cosine distance to clean target-EO representations from a frozen vision foundation model, providing semantic guidance for from-scratch training without annotations or inference-time cost.
-
State-of-the-art results on two benchmarks. Experiments on QXS-SAROPT and SAR2Opt show leading perceptual and distributional performance, with a controlled SD2.1-codec variant isolating the effect of codec selection.
Main Findings
-
Codec choice has a large, measurable effect. Across six pretrained autoencoders and four SAR–EO datasets, FLUX.2 achieves the highest PSNR in six of the eight dataset–modality settings and raises the mean EO reconstruction ceiling by 7.49 dB over SD2.1, the codec used by recent state-of-the-art latent SET methods.
-
Best distributional and perceptual results on QXS-SAROPT. ReFlowSET obtains the best DISTS (0.231) and ties the best FID (19.1) with SD2.1 fine-tuning, at LPIPS 0.534, SSIM 0.355, and PSNR 16.09. The best LPIPS on this benchmark is C-DiffSET's 0.526.
-
Best FID, DISTS, and LPIPS on SAR2Opt. ReFlowSET improves the second-best FID from 71.8 to 66.3 and DISTS from 0.211 to 0.185, and yields the lowest LPIPS (0.522), with SSIM 0.287 and PSNR 16.06. The authors report these as relative gains of 7.7 percent in FID and 12.3 percent in DISTS.
-
The gains come from a from-scratch generator. SD2.1 fine-tuning and C-DiffSET initialize U-Net backbones from pretrained SD2.1 weights, while ReFlowSET trains its DiT from scratch — reaching or surpassing those baselines on perceptual and distributional metrics.
-
Codec swapping is a controlled win. Within the identical ReFlowSET framework, replacing the SD2.1 codec with FLUX.2 improves all five metrics on both benchmarks, reducing FID from 25.5 to 19.1 on QXS-SAROPT and from 84.5 to 66.3 on SAR2Opt.
-
Pixel metrics are not maximized. ReFlowSET does not maximize PSNR or SSIM, which the authors attribute to the inherent ambiguity and local misalignment of paired SAR–EO observations; they emphasize perceptual and distributional metrics while reporting pixel fidelity for completeness.
-
Dual-stream conditioning with channel fusion is the best trade-off. In the ablation on SAR2Opt with 12-block models, dedicated SAR processing with channel fusion reduces FID from 72.865 (shared-stream, input-concatenation baseline) to 70.436, at a parameter cost of 108.46M to 145.04M, so the authors frame it as a system-level architecture comparison.
-
Token-wise fusion is not worth its cost. Token fusion increases training VRAM from 10.63 to 13.57 GB and latency from 1.34 to 2.06 seconds per image while providing virtually no FID benefit (70.450 versus 70.436). ReFlowSET therefore adopts channel-wise fusion.
-
Representation alignment helps for free at inference. Under identical dedicated SAR stream, channel fusion, and 4D+8S topology, adding the alignment loss reduces FID from 71.851 to 70.436, leaves deployed parameters (145.04M) and latency (1.34 seconds) unchanged, and raises peak training VRAM only from 9.80 to 10.63 GB.
-
Final configuration. The deployed ReFlowSET scales the ablation's 4 dual-stream/8 single-stream allocation to 8 and 16 blocks, preserving the same 1:2 topology ratio, within a 24-block, 509.3M-parameter DiT.
Methodology in Plain English
The pipeline starts by choosing the latent space rather than assuming one. For each candidate autoencoder, the authors encode and decode SAR and EO images and measure round-trip PSNR. A high EO round-trip score means the codec does not distort the target the model is asked to reach; a high SAR score means more source structure survives for conditioning. FLUX.2 wins this audit, so its encoder and decoder are frozen and reused for both modalities in all later experiments.
The generator is then trained from scratch inside that latent space. The EO target latent and Gaussian noise are linearly interpolated over a time variable, and a Diffusion Transformer is trained to predict the velocity that moves noise toward the EO latent. The SAR latent is supplied throughout as a persistent condition. Instead of starting the flow from the SAR latent, the model starts from independent noise, which prevents the regression target from becoming a simple analytic function of the inputs.
Conditioning is structured in two stages. For the first blocks, the SAR condition and the noisy-EO latent are projected and processed in separate, independently parameterized streams, preserving modality-specific statistics. Their features are then concatenated along the channel dimension, projected back to model width, and passed through the remaining single-stream blocks that predict the velocity field.
Because the DiT inherits no semantic priors from a pretrained generator, and because flow matching only supervises velocity, a frozen vision foundation model provides additional guidance during training. Clean EO targets are passed through the frozen model to produce reference representations; a trainable projector maps corresponding intermediate noisy-EO DiT features into that same space; and a cosine-distance loss pulls them together. The reference representations use a stop-gradient, and the total objective is the flow-matching loss plus this alignment loss weighted by 0.5, fixed across all experiments. The foundation model and projector are discarded at inference.
At test time the learned ordinary differential equation is integrated from t=0 to 1 and the endpoint is decoded into an EO image. Settings reported include a randomly initialized 24-block, 509.3M-parameter DiT, the frozen FLUX.2 autoencoder, a frozen DINOv3 ViT-L/16 teacher, training on two NVIDIA RTX 4090 GPUs with AdamW at a peak learning rate of 5e-4, a 1k-step warmup, cosine decay, bf16 precision, and 0.1 conditioning dropout, with classifier-free guidance of 1.5 and an Euler solver using 50 function evaluations. QXS-SAROPT is trained for 40k updates at batch 64 and SAR2Opt for 20k updates at batch 32.
Why This Matters
Impact on research. The paper reframes the autoencoder as a design variable in latent translation pipelines and supplies a modality-wise audit protocol that other cross-modal latent methods can reuse. It also shows that a from-scratch generator paired with a strong codec can match or beat baselines that start from pretrained generator weights, and demonstrates that representation alignment from a frozen foundation model can be added to a from-scratch DiT with no inference-time cost — a pattern applicable beyond SAR-to-EO. The ablation additionally shows that where fusion happens matters more than sequence-level detail: token-wise fusion doubled sequence length and latency for essentially no FID improvement over channel fusion.
Real-world applications.
- All-weather Earth observation, producing interpretable optical-style imagery when cloud cover or darkness blocks optical sensors.
- Rapid disaster response, where SAR captures flood, storm, or earthquake conditions and analysts need a visually readable product.
- Maritime and infrastructure monitoring, where synthetic-aperture radar supports vessel and structure detection around the clock.
- Environmental and land-cover analysis, where coherent land-cover appearance in translated imagery aids visual comparison over time.
Industry relevance. The work targets the geospatial and remote sensing industry, where imagery providers and analytics firms need to serve optical-style products from SAR-only acquisitions. The reconstruction audit provides a practical, low-cost way to choose a latent codec before committing training compute, and the alignment branch's zero inference overhead matters for deployment on constrained hardware. Code and pretrained weights are released at the KAIST-VICLab GitHub repository, which lowers the barrier to adoption. Funding is acknowledged from the National Research Foundation of Korea under the Sejong Science Fellowship Program (RS-2026-25484549, 50 percent) and an NRF grant (RS-2025-02222525, 50 percent).
Future Directions
- Cross-sensor generalization. The authors explicitly list examining generalization across different SAR sensors as future work.
- Suppressing unsupported EO-like structures. The conclusion also names suppression of EO-like structures that are not supported by the SAR observation as a direction, addressing hallucination in an ambiguous cross-modal mapping.
- Wider codec and modality audits. The reconstruction-ceiling analysis covered six autoencoders and four datasets; extending the audit to newer codecs and additional sensor pairings is a natural extension.
- More thorough topology and fusion studies. The ablation used compact 12-block models with a 4D+8S split to control cost, while the final model uses 8D+16S; the parameter-matched and depth-scaled behavior of alternative stream allocations and fusion sites remains open.
Target Audience
This paper suits graduate students and researchers working on cross-modal image translation, latent diffusion and flow-matching generative models, and remote sensing image processing. It is also relevant to practitioners in geospatial analytics who need to choose a latent codec or decide whether a from-scratch generator can compete with pretrained-generator fine-tuning, and to engineers interested in training-only representation alignment as a way to add semantic supervision without inference cost. Readers without background in diffusion or flow matching will find the quantitative comparisons readable but the methodology sections demanding.
Authors’ abstract
SAR-to-EO image translation aims to generate electro-optical (EO) imagery from synthetic aperture radar (SAR) observations. Existing latent diffusion approaches typically inherit a predetermined autoencoder, although reconstruction fidelity can vary substantially across codecs and modalities. Because the latent codec affects the round-trip preservation of both SAR conditions and EO targets, codec selection constitutes a fundamental design choice; nevertheless, existing methods largely rely on codecs pretrained on natural images. To remedy this, we introduce ReFlowSET, a conditional latent flow-matching framework that selects its codec through a joint SAR--EO reconstruction audit. Rather than inheriting a heavyweight pretrained generator, ReFlowSET trains a substantially smaller conditional DiT from scratch in the selected latent space, using dual-stream SAR conditioning followed by joint feature refinement. To provide semantic guidance for this from-scratch training, intermediate noisy-EO features are aligned with clean target-EO representations extracted by a frozen vision foundation model. This alignment is used only during training and introduces no additional inference cost. Experiments on QXS-SAROPT and SAR2Opt demonstrate state-of-the-art performance across diverse perceptual fidelity and distributional metrics. Code and pretrained weights are publicly available at https://github.com/KAIST-VICLab/ReFlowSET.