Research
Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Overview Research area: Multimodal image retrieval and agentic vision-language systems, at the intersection of text-to-image search, personal photo-collection management, and LLM-driven search agents.

- arXiv
- 2608.28695
- Published
- 2026-08-27
- Authors
- Rong Shan, Tianyi Xu, Congmin Zheng, Wenteng Chen, Jiachen Zhu, Junjie Wu, Teng Wang, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
AI summary
Overview
- Research area: Multimodal image retrieval and agentic vision-language systems, at the intersection of text-to-image search, personal photo-collection management, and LLM-driven search agents.
- Technical level: Advanced. The paper assumes familiarity with multimodal embedding models (CLIP-style dual encoders, MLLM-based embedders), beam search, and agentic query decomposition, though the core idea is explained in plain language below.
- Scope: The paper defines a new retrieval task called Image Bundle Composition (IBC), releases a benchmark for it (IBCBench), and proposes an agentic framework (BundleWeaver) that composes sets of images satisfying a shared cross-image relation rather than ranking images one at a time.
What This Paper Is About
Standard image search scores every candidate image on its own and returns a ranked list, which cannot express requests like "the transition of the Eiffel Tower from day to night" or "the key moments of a concert night" where meaning lives in the relationship between images. The paper formalizes this gap as Image Bundle Composition: given a text query and a large unstructured photo pool, the system must dynamically assemble a small, cohesive set of images (a bundle) whose members jointly satisfy the query. Because bundles are not predefined or indexed, the search space grows combinatorially, so the authors build a verified benchmark and an agent that grows bundles incrementally instead of enumerating subsets.
Key Contributions
-
A new retrieval paradigm, Image Bundle Composition (IBC). The paper reformulates retrieval from point-wise ranking of isolated images to dynamic composition of cohesive image bundles bound by structural relations such as temporal progression, event summarization, or spatial continuity. It formally shows the search space scales as |𝒮| = Σ_{k=2}^{K_max} (N choose k) ≈ O(N^{K_max}) and that joint relevance Φ(B, q) is non-decomposable, i.e., cannot be recovered by any monotonic aggregation g of individual image scores.
-
IBCBench, the first IBC benchmark. 109,467 images and 667 meticulously verified queries, built through a five-stage semi-automated pipeline anchored to spatiotemporal sessions, with VLM verification and human review against four property criteria (joint completeness, cross-image binding, uniqueness within pool, bounded redundancy). Bundle sizes are distributed as 3 (24.3%), 4 (32.2%), and 5 (43.5%) images, and relations split 52.5% same-location dynamics versus 47.5% cross-location structural ties.
-
BundleWeaver, an agentic framework. It reformulates IBC as query-conditioned incremental hyperedge discovery, combining diverse seed initialization, adaptive sub-query generation by an LLM reasoning agent, contextual candidate pruning, parallel beam search with a composite path score, and whole-bundle VLM reranking.
-
Comprehensive baselines and analysis. The paper establishes four baseline families (multimodal embedding, caption + text embedding, heuristic metadata augmentation, and VLM two-stage decompose-and-rerank), plus component ablations and a backbone-generalizability study across Qwen3-VL-235B, Gemini-3-Flash, Claude Sonnet 4.5, and GPT-4o.
Main Findings
- Atomic matching is fundamentally inadequate for IBC. Multimodal embedding baselines score low even as models scale: CLIP-ViT-B/32 reaches F1 2.21 and EM 0.00, SigLIP2-giant F1 9.97 and EM 0.15, Qwen3-VL-Embed-8B F1 12.14 and EM 0.00, and RzenEmbed F1 14.88 and EM 0.30. Caption + text embedding baselines are similarly low (BM25 F1 3.29, BGE-M3 F1 7.05, Qwen3-Embedding-8B F1 9.32).
- Metadata heuristics do not fix the problem. Session Clustering (F1 4.84), User+Session Clustering (F1 8.94), Time Proximity (F1 6.67), and User+Spatiotemporal (F1 8.80) show that applying spatiotemporal locality as a rigid filter actually degrades F1 relative to plain multimodal embedding, because proximity alone cannot fulfill the relational roles a query demands.
- Static decompose-and-rerank improves coverage but not exactness. The VLM two-stage baseline raises precision and recall but leaves Exact Match extremely low: Qwen2.5-VL-72B (P 11.39, R 10.16, F1 10.67, EM 0.15), GPT-4o (18.50 / 17.34 / 17.84 / 0.60), Gemini-3-Flash (20.89 / 15.53 / 17.06 / 0.60), Qwen3-VL-235B (15.78 / 14.35 / 14.95 / 0.90), and Claude-Sonnet-4.5 (24.79 / 23.01 / 23.74 / 1.50). Independent sub-query retrieval cannot enforce inter-image constraints.
- BundleWeaver leads on every metric. It achieves Precision 30.95, Recall 30.46, F1 30.28, and EM 7.20, the best result in Table 1, corresponding to relative improvements over the strongest baseline of 24.8% (Precision), 32.4% (Recall), 27.5% (F1), and 380.0% (EM).
- Every component contributes. Ablation against the GPT-4o decompose-and-rerank baseline (F1 17.84, EM 0.60): removing Diverse Seed gives F1 27.21 / EM 6.45; removing Contextual Candidate Pruning gives F1 24.69 / EM 5.85; downgrading beam search to greedy expansion gives F1 25.35 / EM 6.30; removing VLM reranking gives F1 26.13 / EM 6.15. Notably, even without pruning the method still far exceeds the vanilla GPT-4o baseline, so the gain is not attributable to metadata pruning alone.
- Larger bundles are harder, but BundleWeaver degrades more slowly. Breakdown analysis by bundle size and location type shows all methods worsen as bundle size increases, while BundleWeaver's drop is smaller, suggesting incremental expansion is more robust than retrieving images independently. RzenEmbed performs notably worse on cross-location bundles, and the decompose-and-rerank baseline improves F1 on both location types but keeps low EM.
- Gains transfer across backbones. With Qwen3-VL-235B, BundleWeaver reaches F1 25.73 / EM 5.25 versus 14.95 / 0.90 for decompose-and-rerank; with Gemini-3-Flash, 25.55 / 4.35 versus 17.06 / 0.60; with Claude Sonnet 4.5, 30.32 / 6.60 versus 23.74 / 1.50; with GPT-4o, 30.28 / 7.20 versus 17.84 / 0.60. The EM surge appears regardless of backbone, indicating the structural design, not a smarter VLM, is what enforces bundle-level constraints.
- Absolute performance remains modest, by the authors' own admission. The impact discussion states that the absolute numbers are modest and that the substantial headroom underscores the intrinsic difficulty of the task, positioning BundleWeaver as a foundational baseline rather than a solved endpoint.
Methodology in Plain English
Building the benchmark. The authors start from YFCC-100M, keeping 57 users' data for a pool of 109,467 photos (all with valid EXIF timestamps, over 98.3% with GPS coordinates), and they deliberately discard YFCC's user-defined album IDs so bundles must emerge from visual and spatiotemporal relations rather than manual groupings. GPT-4o is used to cache captions, word-level tags, and event-level signals, and RzenEmbed embeddings are used to penalize redundancy. Photos are grouped into sessions by splitting wherever there is a time gap over 6 hours or a geographic jump over 20 km; within sessions, sliding windows of size K in [3, 5] are enumerated and scored by a composite heuristic combining spatiotemporal coverage, semantic richness, and a redundancy penalty, with windows below a score of 2.5 discarded. This yields 7,460 diverse candidate windows. Claude-Opus-4.5 then acts as a semantic verifier against predefined relational templates (Same-location Dynamics and Cross-location Structures), generating a natural query and articulating a specific cross-image shared anchor. Finally, four expert annotators review the VLM-approved candidates against the four property criteria, filtering ambiguous cases. The end-to-end acceptance rate is under 9%, producing 667 verified queries.
Composing bundles. BundleWeaver treats the image pool as the vertex set of an implicit, fully connected hypergraph, where a target bundle is a hyperedge connecting a small subset, and the task becomes incremental pathfinding from a seed image. First, an LLM extracts the primary visual anchor from the query and produces an initial search direction, which is used for a global nearest-neighbor search; from the top retrieved images, a diverse set of seeds is chosen so multiple parallel search trees start from different sub-clusters rather than one local optimum. Each branch then expands under a beam search of width B_u. At each step, the LLM reasoning agent looks at the original query plus captions of images already chosen and generates an adaptive sub-query targeting the missing relational role rather than simply the next most similar image. Candidate expansions are restricted to a plausible spatiotemporal neighborhood of the seed to avoid physically disconnected distractors. Surviving branches are selected by a composite score that averages the cosine similarity between each sub-query and its chosen image (step-wise local match) and adds λ times the cosine similarity between the original query and the normalized sum of the selected image embeddings (holistic bundle alignment). Search terminates adaptively when the LLM judges the bundle complete, bounded by K_max. Finally, all completed paths are pooled and a VLM serves as a pointwise reranker, scoring each candidate bundle with its query on structural coherence and cross-image binding on a 1-10 scale; the top-scoring hyperedge is returned.
Evaluation. Because IBC outputs variable-length sets rather than ranked lists, evaluation uses set-level Precision, Recall, and F1 plus Exact Match (the percentage of perfectly predicted ground-truth bundles), averaged per query. For baselines that return ranked lists, an oracle-size evaluation truncates each list to the top-|B*| images, where |B*| is the ground-truth bundle size; since predicted and ground-truth sets then have identical sizes, their Precision, Recall, and F1 become mathematically equivalent, which is why those baseline rows show identical values across the three metrics. BundleWeaver is training-free and zero-shot, built on off-the-shelf foundation models; implementation details are deferred to Appendix D.
Why This Matters
- Impact on research: The paper argues that retrieval systems must evolve from passive, point-wise rankers into active, reasoning-driven constructors capable of inline relational composition. It formalizes non-decomposable joint relevance, showing mathematically why greedy top-K retrieval and query decomposition both fail, and supplies both a benchmark and baseline for others to build on. The work also surfaces what the authors call a critical blind spot in contemporary retrieval methodologies: relational blindness that model scaling alone does not fix.
- Real-world applications:
- Personal photo management, automatically assembling highlights such as the key moments of a concert night or the gradual transition of a landscape from day to dusk from a user's own gallery.
- Trip and travel storytelling, composing a continuous journey spanning multiple landmarks or the day-to-dusk arc of a single landmark.
- Event summarization, turning a large spatiotemporal photo session into a compact, structurally coherent set of images.
- The authors note the technology is envisioned for personal galleries or encrypted local storage, not for surveillance or analysis of third-party subjects.
- Industry relevance: The work is a collaboration involving OPPO alongside Shanghai Jiao Tong University and Shanghai Innovation Institute, and its target use case, searching personal photo collections, maps directly onto phone gallery and cloud photo services. Because BundleWeaver is training-free and works with off-the-shelf backbones (GPT-4o, Claude-Sonnet-4.5, Gemini-3-Flash, Qwen2.5-VL-72B, Qwen3-VL-235B), its approach is deployable without bespoke model training, and the authors report dataset and code are available via GitHub and Hugging Face.
Future Directions
- Extend beyond static personal photos. The limitations section states the current investigation is constrained to static images within YFCC; specialized domains such as medical imaging sequences or legal evidentiary archives, and dynamic modalities such as Video Bundle Retrieval where temporal dynamics are continuous rather than discrete, remain untried.
- Move past training-free, zero-shot operation. BundleWeaver currently relies on off-the-shelf foundation models; the authors identify end-to-end fine-tuning and parameter-efficient tuning methods tailored specifically for IBC as a highly promising direction.
- Close the exact-match gap. With EM at 7.20 even for the best configuration, perfect recovery of a non-decomposable visual narrative from an O(N^{K_max}) combinatorial space remains open, as the authors frame their own method as a starting point rather than an endpoint.
- Improve reasoning over large bundles and cross-location structures. Performance degrades for all methods as bundle size grows and remains weakest on cross-location bundles where visually diverse images are hard to connect through point-wise similarity, suggesting headroom in how relational roles are modeled across greater visual and spatial dispersion.
Target Audience
This paper is most useful to researchers and engineers working on multimodal retrieval, image search infrastructure, and agentic LLM/VLM systems, particularly those building gallery search or photo-organization products at consumer scale. It also suits benchmark builders and evaluation specialists interested in set-level metrics for non-decomposable retrieval targets, and graduate students looking for a clearly formalized new task with released data, code, and a strong baseline to improve on. Readers should have some grounding in embedding-based retrieval and LLM agent design to fully follow the methodology and ablation reasoning.
Authors’ abstract
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural relations rather than isolated snapshots. To address this limitation, we introduce **Image Bundle Composition (IBC)**, a novel paradigm that shifts the objective from ranking individual images to dynamically composing cohesive image bundles from a massive, unstructured photo pool. Since target bundles are not predefined, IBC presents a severe combinatorial explosion challenge and demands modeling non-decomposable joint relevance. To establish this paradigm, we construct **IBCBench**, the first IBC benchmark dataset containing 109,467 images and 667 verified queries, built via a semi-automated verification pipeline. Furthermore, we propose **BundleWeaver**, an agentic framework that reformulates IBC as query-conditioned incremental hyperedge discovery. By employing a Large Language Model to adaptively search for missing relational roles and utilizing a Vision-Language Model for whole-bundle verification, BundleWeaver effectively navigates the combinatorial space. Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition. Our dataset and code are available.