Skip to content
AI.info

Research

Match-and-Fuse: Consistent Generation from Unstructured Image Sets

Overview Research area: Computer vision / generative AI — controllable text-to-image generation applied to sets of images rather than single images. Technical level: Advanced. Familiarity with diffusi

Match-and-Fuse: Consistent Generation from Unstructured Image Sets
arXiv
2511.22287
Published
2025-11-27
Authors
Kate Feingold, Omri Kaduri, Tali Dekel

AI summary

Overview

Research area: Computer vision / generative AI — controllable text-to-image generation applied to sets of images rather than single images.

Technical level: Advanced. Familiarity with diffusion models, attention/feature manipulation (keys, values, RoPE), and correspondence estimation is needed to follow the method details. The high-level idea is accessible to a general reader; the mechanics are not.

Scope (one sentence): The paper introduces a zero-shot, training-free framework that jointly generates a whole set of consistent images from an unstructured source set, using a pairwise graph of image-grid generations combined with match-guided feature fusion.

What This Paper Is About

Most generative image tools work on one image at a time, or on dense video sequences that provide strong temporal cues. Image sets — photo albums, product catalogs, real-estate listings, storyboards — sit awkwardly between the two: they show shared content (the same object, character, or product) but from different viewpoints, times, poses, and backgrounds, with no temporal continuity and no stable 3D structure to exploit.

The paper's goal is set-to-set generation: given a source set of images plus user prompts describing the desired shared content and overall theme, produce a new set that keeps the shared elements visually and semantically consistent across all images, while leaving non-shared regions (backgrounds, etc.) free to vary.

Key Contributions

  1. The first method for unstructured set-to-set generation — it moves beyond image-pair editing to entire collections, handling sets of varying sizes (the paper reports working ranges of 2–15 images, and up to 20 images under the full method within memory constraints).
  2. A flexible, automated, training-free, mask-free framework — it requires only simple text prompts as input and operates zero-shot on a frozen pre-trained text-to-image model, without fine-tuning, masks, or manual supervision.
  3. A graph-based formulation (the Pairwise Consistency Graph) that consolidates all pairwise generations into one joint framework, enforcing local consistency between image pairs and global coherence across the set, while scaling linearly in runtime by limiting each node to 4 random neighbors.
  4. A new consistency metric, DINO-MatchSim, for fine-grained cross-image consistency, which the authors report aligns more strongly with human judgments than global similarity metrics.

Main Findings

  • Highest reported consistency and visual quality among compared methods. On the paper's benchmark, Match-and-Fuse scores DINO-MatchSim 0.80 ± 0.003 and DreamSim 0.85 ± 0.004, versus Edicho (0.72 ± 0.004 / 0.81 ± 0.004), FLUX (0.66 ± 0.004 / 0.76 ± 0.004), IC-LoRA (0.65 ± 0.004 / 0.71 ± 0.006), and FLUX Kontext (0.57 ± 0.004 / 0.78 ± 0.004). Its CLIP score is 0.66 ± 0.001, comparable to the baselines (range 0.65–0.67).
  • Human and VLM preference favor the method over every baseline. In 2AFC comparisons, users preferred Match-and-Fuse over FLUX Kontext in 88% of cases, over IC-LoRA in 90%, over FLUX in 92%, and over Edicho in 83%; the GPT-5-based VLM evaluation gave 82%, 92%, 94%, and 78% respectively.
  • The new metric agrees with humans better than prior metrics. Agreement with the users' majority vote was 91.4% for DINO-MatchSim, versus 84.9% for the VLM and 84.3% for DreamSim.
  • Every component contributes to consistency. Removing Feature Guidance drops DINO-MatchSim to 0.76 ± 0.004; removing Multiview Feature Fusion drops it to 0.78 ± 0.003; removing the Pairwise Consistency Graph (falling back to single-image predictions) drops it to 0.75 ± 0.004 — all below the full method's 0.80 ± 0.003.
  • Feature similarity tracks visual consistency. Across progressively more consistent settings (random images → descriptive prompts → control signals → grid generation → DDIM inversion), cosine similarity at matched locations rises, based on measurements averaged over 40 image pairs.
  • The method degrades with set size but stays stronger than baselines. It remains more consistent for 9 images than the baselines are for 2; a fully connected graph variant performs only slightly better than the degree-4 sparse graph used in all experiments.
  • Robustness under sparse matches. DINO-MatchSim stays high even when a large percentage of correspondences is randomly removed, whereas the variant without Feature Guidance degrades more rapidly.
  • Runtime scales linearly for N ≥ 5, with per-image runtime approximately constant from that point on; RoMA matching costs 198.8 ms per image pair on an RTX6000 GPU. Measurements were run on an NVIDIA A100 GPU at 512 × 512 resolution.

Methodology in Plain English

The method starts from a frozen, pre-trained depth-conditioned text-to-image diffusion model (FLUX with depth conditioning) that is known to produce multi-image grids when prompted with layout-style prompts — the authors call this the grid prior. This prior gives partial cross-image consistency, but it degrades as more images are crammed onto one canvas and it is bounded by the model's native resolution.

To work around this, the authors model the image set as a graph: each image is a node, and each edge is a two-image grid generation. Every denoising step, the noisy latents for all images are assembled into these pairwise grids, denoised jointly, and then the multiple versions of each image produced by its different edges are extracted and averaged back into a single per-image latent. This lets the method exploit the grid prior without inheriting its resolution limit.

Two mechanisms enforce consistency:

  • Multiview Feature Fusion (MFF): using dense 2D correspondences computed from the source images (RoMA matches filtered by a confidence threshold of c > 0.05), the method averages internal keys and values at matched locations across images and across all graph edges. This is applied to keys and values before RoPE.
  • Feature Guidance: a lighter, gradient-based refinement that minimizes the distance between feature maps at matched locations via the latents, correcting residual fine-grained misalignment. It is applied only where the paper's schedule specifies, as a light touch rather than the primary driver.

Prompts are composed automatically: a vision-language model (GPT-4o) writes per-image captions describing non-shared content while the user supplies the shared-content and theme prompts. The method needs no object masks — correspondences alone identify shared regions.

The pipeline is evaluated on a purpose-built benchmark of 400 edits across 149 distinct image sets of 3–15 images, combining subject-driven sets, 3D datasets, video keyframes, and ChatGPT-generated sketch storyboards.

Why This Matters

Research impact. The paper opens an under-explored modality — collections of unordered images — and provides both a formulation and an evaluation tool. The graph-based consolidation trick and the finding that matched-feature similarity predicts visual consistency are transferable ideas. The authors also report negative results on adapting UNet-style "extended attention" to a DiT (FLUX), documenting that reusing, extending, or match-warping positional embeddings leads to duplicated generations, artifacts, or a severe quality–consistency tradeoff — useful information for follow-up work.

Real-world applications (as described in the paper):

  • Product advertising and catalogs — transforming fixed multi-view product layouts into coherent themed edits, including "multi-cut" ad edits with cross-cut consistency that one-to-one editing cannot achieve.
  • Character concept art and storyboards — consistent visualization of hand-drawn sketches across a set of frames.
  • Film set design and creative workflows — applying a shared theme across a multi-view layout while keeping the shared subject coherent.
  • Localized editing — background-preserving edits by combining the method with FlowEdit, demonstrated in the paper's extended applications.

Industry relevance. The approach is training-free and requires no fine-tuning, LoRA training, or mask annotation, which lowers the barrier to integration into existing creative tooling and content-production pipelines. Its runtime stays roughly constant per image once sets reach N ≥ 5, though the authors cap generation at 20 images with the full method (22 without Feature Guidance) due to memory.

Future Directions

  • Scaling to larger and different modalities. The authors name video collections and foundation models for set-to-set generation as natural extensions, and note memory caps generation at 20 images with the full method or 22 without Feature Guidance.
  • Reducing dependence on correspondence quality. Performance depends on the quality and density of pixel correspondences; largely unmatched or ambiguous regions (disocclusions, symmetries) can produce inconsistencies.
  • Improving base-model fidelity to source structure. The method relies on the base model preserving the source conditioning depth maps, which occasionally deviates from the source layout depending on the prompt and the model's generative prior (the authors cite body and hand poses as an example).
  • Automating localized-editing trade-offs. With FlowEdit integration, balancing structure preservation against appearance change often still requires per-edit hyperparameter tuning; automating that selection is left to future work.

Target Audience

Researchers and graduate students working on controllable image generation, image editing, and multi-image consistency will get the most from this paper, particularly those interested in training-free methods that manipulate diffusion internals. Practitioners building creative tools — advertising, e-commerce catalog production, concept art, storyboarding — will also find the capabilities and the reported runtime/memory limits directly relevant. Readers without a diffusion-model background can still follow the graph formulation and the empirical comparisons, but the feature-fusion details assume some familiarity with transformer attention and positional embeddings.

Authors’ abstract

We present Match-and-Fuse - a zero-shot, training-free method for consistent controlled generation of unstructured image sets - collections that share a common visual element, yet differ in viewpoint, time of capture, and surrounding content. Unlike existing methods that operate on individual images or densely sampled videos, our framework performs set-to-set generation: given a source set and user prompts, it produces a new set that preserves cross-image consistency of shared content. Our key idea is to model the task as a graph, where each node corresponds to an image and each edge triggers a joint generation of image pairs. This formulation consolidates all pairwise generations into a unified framework, enforcing local consistency while ensuring global coherence across the entire set. This is achieved by fusing internal features across image pairs, guided by dense input correspondences, without requiring masks or manual supervision, and by leveraging an emergent prior in text-to-image models that encourages coherent generation when multiple views share a single canvas. Match-and-Fuse achieves state-of-the-art consistency and visual quality, and unlocks new capabilities for content creation from image collections.

Read the original paper