Research
Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
Overview Research area: Computer vision — controllable image generation for remote sensing semantic segmentation, combining Diffusion Transformers (DiTs) with flow-matching sampling. Technical level:

- arXiv
- 2512.16740
- Published
- 2025-12-18
- Authors
- Yunkai Yang, Yudong Zhang, Kunquan Zhang, Jinxiao Zhang, Xinying Chen, Haohuan Fu, Runmin Dong
AI summary
Overview
Research area: Computer vision — controllable image generation for remote sensing semantic segmentation, combining Diffusion Transformers (DiTs) with flow-matching sampling.
Technical level: Advanced. The paper assumes familiarity with diffusion/flow-based generative models, attention-based multimodal architectures, and standard semantic segmentation evaluation.
Scope: The paper proposes TODSynth, a training-data synthesis framework that pairs a mask-conditioned MM-DiT generator with a task-feedback-guided sampling method (CRFM), and evaluates whether the resulting synthetic remote sensing imagery improves downstream segmentation.
What This Paper Is About
Building pixel-level labeled datasets for remote sensing semantic segmentation is expensive and slow, so researchers have turned to controllable generative models that synthesize images from existing semantic masks. However, two problems limit the usefulness of this synthetic data: it is unclear which mask-conditioning scheme works best for DiT-based generators, and the stochastic sampling process often produces images that drift away from the semantic mask they were supposed to follow. The paper addresses both by systematically comparing control schemes and introducing a sampling-time correction guided by a downstream segmentation model's loss.
Key Contributions
-
TODSynth framework. A task-oriented data synthesis framework that combines architecture-level control (a Multimodal Diffusion Transformer with unified triple attention) and sampling-level control to improve synthetic data for remote sensing semantic segmentation.
-
Control-Rectify Flow Matching (CRFM). A plug-and-play sampling strategy that uses a pre-trained segmentation network's cross-entropy loss to compute a rectification vector applied to the predicted velocity field during the early, high-plasticity stage of sampling—without retraining the diffusion model.
-
Systematic evaluation of mask-to-image control schemes. The authors compare unified triple attention, siamese MM-attention, and mask-adapter conditioning, showing that text–image–mask joint attention with full fine-tuning of image and mask branches performs best.
-
Extensive experiments on two remote sensing benchmarks. TODSynth outperforms state-of-the-art controllable generation methods on FUSU-4k and LoveDA-5k, and the code is released at https://github.com/Yunkai-Yang/crfm.
Main Findings
-
Tri-attention beats the alternatives. On FUSU-4k, unified triple attention reaches OA 75.41, mIoU 48.57, mAcc 61.67, versus siamese MM-attention (74.94 / 48.46 / 61.44), mask-adapter (74.94 / 47.41 / 59.62), and ControlNet with SD v1.5 (73.85 / 45.13 / 56.77).
-
TODSynth improves over a fully supervised real-data baseline. Gains of 1.39% / 4.14% / 6.83% in OA/mIoU/mAcc on FUSU-4k, and 1.60% / 2.08% / 2.22% on LoveDA-5k, relative to the baseline trained only on real data.
-
CRFM contributes independently of the architecture. Compared against SD v3.5 with the identical architecture and post-processing filter, adding CRFM improves mIoU/mAcc by 0.84% / 1.6% on FUSU-4k and 0.67% / 1.02% on LoveDA.
-
Fewer synthetic samples are needed. TODSynth achieves better results using a synthetic-to-real ratio of 3, while comparison methods use ratios of 5 or 10.
-
Pixel-level filtering outperforms image-level filtering. On FUSU-4k, the FreeMask pixel-level filter gives 75.41 / 48.57 / 61.67, compared with 74.78 / 46.69 / 58.08 for a CLIP-score image-level filter; adding CRFM on top of pixel-level filtering gives 75.66 / 49.41 / 63.27.
-
CRFM is sensitive to how many steps it is applied. With 23 sampling steps, applying CRFM for 4 steps gives the best downstream result (mIoU 49.41, FID 38.65), while 6 steps degrades FID to 66.95. The same pattern appears at 18 steps (FID rises to 86.69 at 6 CRFM steps) and 13 steps (FID 132.92 at 6 CRFM steps).
-
DiT-based generators control remote sensing masks better than UNet-based ones. This is stated to be particularly evident on FUSU, which contains 17 diverse land-cover categories, though the model comparison includes SynthEarth as a remote sensing generative foundation model used without tuning.
-
Optimizing the velocity field instead of the latent avoids mode collapse. The authors report that directly optimizing the latent state with the downstream task gradient collapses to a limited subset of samples, whereas rectifying the velocity field produces stable, continuous corrections.
Methodology in Plain English
The pipeline has three stages.
First, a generator is trained. The authors start from the pre-trained Stable Diffusion 3.5 model, which uses a Multimodal Diffusion Transformer architecture, and add a third input stream for the semantic mask alongside text and image latents. Rather than routing mask information through a separate adapter, all three streams are concatenated into a single attention computation, so text semantics and spatial mask constraints interact directly. The image and mask branches are fully fine-tuned. Text prompts carry global semantic information, which compensates for the lack of fine-grained descriptions in remote sensing scenes.
Second, images are sampled with a correction mechanism. At each of the earliest sampling steps, CRFM decodes the model's current prediction of the final image, feeds it to an already-trained segmentation network, and computes cross-entropy loss against the ground-truth mask. The gradient of that loss with respect to the predicted velocity is used as a rectification vector, scaled by a tunable factor and added to the velocity. This nudges the sampling trajectory toward mask-consistent output while the trajectory is still flexible. The correction is deliberately applied only in early steps, because late corrections amplify the segmentation model's own prediction errors and produce adversarial-looking artifacts. Crucially, the correction modifies the velocity field rather than the latent state directly, which the authors say avoids mode collapse.
Third, the corrected synthetic images are filtered—either by a CLIP-score rule retaining image–object–background triplets or by the FreeMask pixel-level filter—and combined with real images to train downstream segmentation models.
Training used 512×512 resolution, the AdamW optimizer with a learning rate of 1×10⁻⁵ and weight decay of 0.01, 200,000 steps, and 8 NVIDIA RTX 4090 GPUs. Downstream segmentation used the MMSegmentation framework with standard augmentations including random crop, random flip, and PhotoMetricDistortion.
Why This Matters
Impact on research. The paper shifts attention from post-hoc filtering of synthetic data to in-process control during sampling, and it argues that the right architectural choice of mask-conditioning matters more than was previously established for DiT-based remote sensing generation. It also provides evidence that correcting velocity fields, rather than latent states, is the more stable optimization target for task-guided generation.
Real-world applications:
- Land-use and land-cover classification over large regions where manual pixel labeling is prohibitively expensive.
- Environmental monitoring and change detection, where rare land-cover categories are underrepresented in existing datasets.
- Thematic mapping at multiple spatial resolutions, where labeled samples are scarce in specific regions or sensor types.
- Rapid dataset bootstrapping for new sensors or geographic areas with no existing annotations.
Industry relevance. Organizations that depend on Earth observation—agriculture, urban planning, forestry, disaster response—frequently lack labeled imagery for their specific target regions. A data synthesis method that improves downstream accuracy using fewer synthetic samples and a synthetic-to-real ratio of 3 could reduce both annotation budgets and generation compute. The method also requires no retraining of the diffusion backbone, which lowers adoption cost.
Future Directions
- Robustness of CRFM to the pre-trained segmentation model. The authors note CRFM needs a sufficiently capable segmentation network to give effective guidance, and that with very limited data it loses effectiveness. They propose exploring remote sensing foundation models as the guide.
- Diversity and cross-domain performance. The current work does not use a strong remote sensing generative foundation model or external remote sensing reference data, which the authors identify as limiting synthetic sample diversity.
- Automatic selection of CRFM step count. Performance peaks at a particular number of correction steps and degrades beyond it, so a principled way to choose this hyperparameter remains open.
- Tension between downstream accuracy and FID. At high CRFM step counts, downstream segmentation performance can still rise while FID deteriorates sharply, raising the question of how to measure synthetic data quality when the two metrics disagree.
Target Audience
Researchers and practitioners working on remote sensing semantic segmentation, controllable image generation, and synthetic training data. It is most useful to readers with a working knowledge of diffusion or flow-matching models and attention architectures, and to applied teams looking to augment scarce pixel-level annotations with generated imagery.
Authors’ abstract
With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semantic mask control and the uncertainty of sampling quality often limit the utility of synthetic data in downstream semantic segmentation tasks. To address these challenges, we propose a task-oriented data synthesis framework (TODSynth), including a Multimodal Diffusion Transformer (MM-DiT) with unified triple attention and a plug-and-play sampling strategy guided by task feedback. Built upon the powerful DiT-based generative foundation model, we systematically evaluate different control schemes, showing that a text-image-mask joint attention scheme combined with full fine-tuning of the image and mask branches significantly enhances the effectiveness of RS semantic segmentation data synthesis, particularly in few-shot and complex-scene scenarios. Furthermore, we propose a control-rectify flow matching (CRFM) method, which dynamically adjusts sampling directions guided by semantic loss during the early high-plasticity stage, mitigating the instability of generated images and bridging the gap between synthetic data and downstream segmentation tasks. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art controllable generation methods, producing more stable and task-oriented synthetic data for RS semantic segmentation.