Research
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval Overview Research area: Computer Vision — specifically Composed Image Retrieval (CIR) using foundation models (Large Multimod

- arXiv
- 2602.00813
- Published
- 2026-01-31
- Authors
- Tong Wang, Yunhan Zhao, Shu Kong
AI summary
Generating a Paracosm for Training-Free Zero-Shot Composed Image RetrievalOverview
Research area: Computer Vision — specifically Composed Image Retrieval (CIR) using foundation models (Large Multimodal Models, Vision-Language Models, and text-to-image generators).
Technical level: Advanced. The paper assumes familiarity with retrieval metrics (Recall@k, mAP@k), vision-language encoders (CLIP/OpenCLIP), and multimodal generation pipelines.
Scope: A single paper proposing a training-free zero-shot CIR method, Paracosm, that generates a "mental image" per query plus synthetic counterparts of database images to improve retrieval.
What This Paper Is About
Composed Image Retrieval lets a user search a database with a reference image plus a text instruction describing how to change that image. The problem is that the desired result — the "mental image" the user has in mind — is never physically available; only the query implicitly defines it. Prevailing zero-shot approaches sidestep this by having a Large Multimodal Model write a textual description of the intended target, then doing text-to-image matching. This paper instead has the model generate the mental image directly, and matches images to images.
Key Contributions
-
Solving CIR from first principles via mental image generation. Paracosm prompts an LMM to generate a "mental image" for each multimodal query (by editing the reference image according to the modification text), enabling image-to-image matching instead of relying solely on textual descriptions.
-
Mitigating the synthetic-to-real domain gap. Because generated mental images are synthetic, the authors generate a synthetic counterpart for every real image in the database and match both the real photo and its synthetic counterpart against the query. This is a departure from prior ZS-CIR work, which the paper states has not considered generating descriptions or synthetic images for database images.
-
A training-free zero-shot method. Paracosm trains no new models; it uses pretrained LMMs (Qwen2.5-VL-7B-Instruct, Qwen-Image, Qwen-Image-Edit) and pretrained VLMs for feature extraction and matching.
-
State-of-the-art zero-shot results. The method outperforms existing zero-shot CIR methods across CIRR, CIRCO, and Fashion IQ benchmarks with both ViT-L/14 and ViT-G/14 backbones, and rivals some supervised methods.
Main Findings
-
CIRR, ViT-L/14 backbone: Paracosm reaches R@1 = 31.95, R@5 = 61.56, R@10 = 72.96, with Recall_Subset@1 = 64.68, @2 = 82.89, @3 = 91.47. The strongest training-free comparison reported, OSrCIR, is listed at R@1 = 37.26 under ViT-G/14, so the backbone matters substantially; under ViT-L/14 the best non-Paracosm zero-shot entries include LDRE (R@1 = 26.53) and IP-CIR + LDRE (R@1 = 29.76).
-
CIRR, ViT-G/14 backbone: Paracosm R@1 = 39.30, R@5 = 70.41, R@10 = 80.39, Recall_Subset@1 = 70.82, @2 = 86.92, @3 = 94.46. OSrCIR is reported at R@1 = 37.26 and CoTMR at 36.36.
-
CIRCO, ViT-L/14: mAP@5 = 30.24, mAP@10 = 31.51, mAP@25 = 34.29, mAP@50 = 35.42, versus LDRE at mAP@5 = 23.35 and IP-CIR + LDRE at 26.43.
-
CIRCO, ViT-G/14: mAP@5 = 39.82, mAP@10 = 40.86, mAP@25 = 43.96, mAP@50 = 45.05, versus CoTMR at mAP@5 = 32.23 and OSrCIR at 30.47.
-
Fashion IQ validation set, ViT-G/14: Shirt R@10 = 40.48 / R@50 = 57.80; Dress 33.17 / 55.18; Toptee 42.58 / 64.20; Average 38.74 / 59.06. The best listed comparison, OSrCIR, averages 37.57 / 57.11.
-
Fashion IQ validation set, ViT-L/14: Shirt 31.80 / 49.51; Dress 24.99 / 47.45; Toptee 31.82 / 52.83; Average 29.45 / 49.93.
-
Image editing beats text-to-image generation for mental images. On the CIRR test set with CLIP ViT-B/32, T2I Generation gives R@1 = 31.71, R@5 = 61.37, R@10 = 73.59, R@50 = 91.54; Image Edit with Qwen gives 32.27 / 62.60 / 75.16 / 92.60; Image Edit with LongCat gives 32.12 / 62.20 / 74.43 / 92.24.
-
Stable across image generators. Swapping Qwen-Image-Edit for LongCat-Image-Edit without tuning hyperparameters or prompts produces stable performance (numbers above).
-
The modification-text weight λ = 0.3 is consistently best. The hyperparameter was tuned on the CIRR validation set using CLIP ViT-B/32, and the paper states λ = 0.3 yields the highest numeric metrics on all benchmarks.
-
Ablation confirms all three ingredients help. With CLIP ViT-B/32 on CIRR/CIRCO, the full configuration (text query + mental image + modification text + database images + synthetic counterparts) achieves CIRR R@1 = 32.27, R@5 = 62.60, R@10 = 75.16, R@50 = 92.60, and CIRCO mAP@5 = 26.10, @10 = 27.02, @25 = 29.29, @50 = 30.45. Removing components degrades results — for example, dropping synthetic counterparts yields CIRR R@1 = 27.93 and CIRCO mAP@5 = 18.29.
-
Higher computational cost, but offloaded offline. On CIRCO (over 123K database images), Paracosm requires 12.9 hours of one-time preprocessing, 41.4 GB of image storage (not retained permanently — only 0.38 GB of compact features), 2.7 GB inference memory, and 14 seconds per inference, with CIRCO mAP@5 = 37.40, CIRR R@1 = 38.24, Fashion IQ R@10 = 36.45 under OpenCLIP ViT-L/14. Comparable methods listed use 0.1 hours preprocessing, 0.38 GB features, and inference of 24 seconds (AutoCIR), 1 second (CoTMR), and 3 seconds (OSrCIR), with lower accuracy.
-
A reproducibility concern is raised. The authors report they could not reproduce OSrCIR's published numbers using CLIP as described, but found that using OpenCLIP reproduces them closely (for example, OSrCIR with OpenCLIP ViT-L/14 and GPT-4o yields CIRCO mAP@5 = 22.75 and Fashion IQ R@10 = 31.55, close to the published 23.87 and 33.26). Paracosm with Qwen2.5-VL under OpenCLIP ViT-L/14 reaches mAP@5 = 37.40 and R@10 = 36.45.
Methodology in Plain English
-
Understand the query. Given a reference image plus a modification text, an LMM with image-editing ability (Qwen-Image-Edit in the final design) edits the reference image according to the text, producing a "mental image" of what the user wants.
-
Describe the mental image. The same LMM is prompted to write a single-sentence description of that mental image, focusing only on visual content and minimizing aesthetic details.
-
Prepare the database. Every real image in the database is passed through an LMM (Qwen2.5-VL-7B-Instruct) to produce a detailed description covering all visible objects, attributes, spatial relationships, and fine-grained visual elements. That description becomes the prompt for a text-to-image model (Qwen-Image) that generates a synthetic counterpart of the database image. This is done offline, ahead of time.
-
Build features. A pretrained VLM supplies a visual encoder V(·) and text encoder T(·). The query feature combines the mental image embedding, the short description embedding, and the modification text embedding, weighted by λ (set to 0.3). Each database image feature is the sum of the real image embedding and its synthetic counterpart's embedding. Images are generated at 512×512 resolution with default parameters from the official implementations.
-
Match. Cosine similarity between the query feature and each database feature, with the highest-scoring image returned as the target.
The overall idea is that both the query and the database are lifted into the same synthetic "paracosm," so synthetic images are compared with synthetic images rather than synthetic against real. Database preprocessing used a cluster of 16 NVIDIA A100 GPUs; individual methods ran on a single A100.
Why This Matters
Impact on research. The paper reframes zero-shot CIR away from "generate a caption, then do text-to-image retrieval" toward "generate the image the user is imagining." It also supplies a methodological caution for the field: differences in VLM backbone choice (CLIP vs. OpenCLIP) can account for large, previously unexplained gaps between reported and reproducible results, which the authors link to community questions raised on the OSrCIR GitHub repository.
Real-world applications:
- E-commerce search — starting from a photo of a garment and searching an online shop for a different genre or style specified in text.
- Fashion retrieval — the paper cites the fashion industry explicitly as a setting where users want to alter a clothing photo's style.
- Personalized visual intelligence — the abstract frames CIR itself as a step toward personalizing visual search.
- General web search — CIR extends traditional text-to-image and image-to-image retrieval tasks used in search services.
Industry relevance. The paper argues CIR remains necessary even as image generation improves, because a generated mental image cannot replace a real product that actually exists in inventory. It also argues the cost profile is acceptable: database processing is a one-time offline expense, and inference-time overhead is comparable to an existing method (14 seconds versus 24 seconds for AutoCIR, though slower than CoTMR's 1 second and OSrCIR's 3 seconds).
Future Directions
- Improving generative fidelity. The authors state Paracosm's performance is inherently limited by the LMMs used, and that generated images can look plausible but lack factual fidelity and fine-grained details (they give the example of a mental image containing a cartoon-style duck that does not correspond to the intended content).
- Reducing computation cost. The paper acknowledges high cost for generating mental images and synthetic database counterparts, and points toward efficient inference via model optimization and optimized implementations as the long-term remedy.
- Safety and content filtering. The authors note Paracosm has no alerting mechanism for inappropriate or malicious multimodal queries, leaving filtering of queries and retrieved targets as open work.
- Reducing reliance on generated intermediates. Because performance hinges on generated descriptions and images, an open question is how much of the pipeline could avoid them or correct their factual errors.
Target Audience
Researchers and graduate students working on multimodal retrieval, zero-shot or training-free methods, and foundation-model pipelines will get the most from this paper. It is also relevant to practitioners building image search for e-commerce and fashion, and to anyone concerned with benchmark reproducibility in the CIR literature. Readers need prior familiarity with CLIP-style encoders and retrieval metrics; beginners will find the equations and model stack demanding.
Authors’ abstract
Composed Image Retrieval (CIR) is the task of retrieving a target image from a database using a multimodal query, which consists of a reference image and a modification text. The text specifies how to alter the reference image to form a ''mental image'', based on which CIR should find the target image in the database. The fundamental challenge of CIR is that this ''mental image'' is not physically available and is only implicitly defined by the query. The contemporary literature pursues zero-shot methods and uses a Large Multimodal Model (LMM) to generate a textual description for a given multimodal query, and then employs a Vision-Language Model (VLM) for textual-visual matching to search for the target image. In contrast, we address CIR from first principles by directly generating the ''mental image'' for more accurate matching. Particularly, we prompt an LMM to generate a ''mental image'' for a given multimodal query and propose to use this ''mental image'' to search for the target image. As the ''mental image'' has a synthetic-to-real domain gap with real images, we also generate a synthetic counterpart for each real image in the database to facilitate matching. In this sense, our method uses LMM to construct a ``paracosm'', where it matches the multimodal query and database images. Hence, we call this method Paracosm. Notably, Paracosm is a training-free zero-shot CIR method. It significantly outperforms existing zero-shot methods on challenging benchmarks, achieving state-of-the-art performance for zero-shot CIR.