Skip to content
AI.info

Research

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing Overview Research area: Computer vision and multimodal machine learning, specifically instruction-based (text-guided) image editin

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
arXiv
2510.19808
Published
2025-10-22
Authors
Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, Zhe Gan

AI summary

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Overview

Research area: Computer vision and multimodal machine learning, specifically instruction-based (text-guided) image editing and the training data that supports it.

Technical level: Intermediate. The paper is a dataset-and-pipeline description rather than a new model architecture, so the core ideas are accessible, but familiarity with supervised fine-tuning, preference alignment (DPO, reward modeling), and multimodal model evaluation helps.

Scope (one sentence): The paper introduces Pico-Banana-400K, a publicly released dataset of approximately 400K text-guided image edits built from real OpenImages photographs, generated with Nano-Banana (Gemini-2.5-Flash-Image), quality-filtered by Gemini-2.5-Pro, and organized under a 35-type editing taxonomy with additional multi-turn and preference subsets.

What This Paper Is About

Progress on text-guided image editing is limited because open research lacks large-scale, high-quality, fully shareable editing datasets built from real images; existing datasets often come from proprietary synthetic generations or small human-curated subsets, and suffer from domain shift, unbalanced edit-type distributions, and inconsistent quality control. The authors address this by constructing and releasing Pico-Banana-400K, roughly 400K instruction-based edit pairs derived from real OpenImages photographs, with automated quality scoring, a systematic edit taxonomy, and specialized subsets for alignment and multi-turn research. The goal is to provide a single, openly shareable corpus rich enough to train and benchmark the next generation of editing models.

Key Contributions

  1. A large-scale, shareable editing dataset. Pico-Banana-400K contains approximately 400K high-quality image editing examples built from real images, organized by a 35-type editing taxonomy across 8 major categories, with quality control through automated scoring plus manual verification. The pipeline is reported to cost approximately 100K USD, and code is released at github.com/apple/pico-banana-400k.

  2. Multi-objective training support. Beyond 258K single-turn supervised fine-tuning examples, the release includes 56K preference pairs (successful vs. failed edits) intended for alignment methods such as DPO and reward modeling, enabling research on robustness and preference learning.

  3. Support for complex editing scenarios. The dataset includes 72K multi-turn editing sequences, each session containing 2–5 consecutive edits, to support research on iterative refinement, context-aware editing, and editing planning.

  4. Dual instruction formats. Every example can carry two parallel instruction variants: a long, detailed instruction from Gemini-2.5-Flash optimized for training, and a short, user-style instruction produced by Qwen2.5-7B-Instruct using human-written annotations as in-context demonstrations.

Main Findings

  • Dataset composition. Figure 1 reports 386K examples split across single-turn SFT (66.8%), preference pairs (14.5%), and multi-turn sequences (18.7%). The abstract and introduction describe the resource as a 400K-image dataset with roughly 400K text-guided edits.

  • Quality gating works via an automated judge. Gemini-2.5-Pro scores each edit on four weighted criteria: Instruction Compliance (40%), Seamlessness (25%), Preservation Balance (20%), and Technical Quality (15%). Edits scoring above a strict threshold empirically set to approximately 0.7 are labeled successful; those below are labeled failures. Successful edits number approximately 258K and failure cases approximately 56K.

  • Failed attempts are reused, not discarded. Up to three retries are allowed per (image, instruction) pair. If a successful edit is reached after one or two failed attempts, the negative edits are retained to form the preference data; if all three attempts fail, the case is discarded from the released set.

  • Global and stylistic edits are the most reliable. Strong artistic style transfer achieves a success rate of 0.9340, film grain/vintage 0.9068, and modern↔historical restyling 0.8875. These operations reshape global texture, color statistics, and tone, requiring limited spatial reasoning.

  • Object-semantic and scene-level edits are moderately reliable. Remove object reaches 0.8328, replace category 0.8348, seasonal change 0.8015, and photo→cartoon/sketch 0.8006. Typical failures involve imperfect localization under text-only conditioning and modest color or texture drift.

  • Precise geometry, layout, and typography are the hardest. Relocate object is the most difficult at 0.5923, change size/shape/orientation attains 0.6627, and outpainting 0.6634 (boundary continuity problems). Text operations are brittle: change font/style yields the lowest rate at 0.5759, while translate/replace/add text remain unstable. Among human stylizations, Pixar/Disney-like 3D reaches 0.6463 and caricature 0.5884, with identity drift and shading artifacts under large shape exaggeration.

  • Some edit types were deliberately excluded. Brightness/contrast/saturation adjustment and sharpen/blur were dropped because changes were negligible or unstable; edits that strongly rewrite an object's perspective or pose were dropped as prone to structural artifacts; and two-image composition (merging objects from two different inputs) was judged not reliable enough for training pairs.

  • Multi-turn construction specifics. 100K single-turn examples were uniformly sampled, and for each one 1–4 additional edit types were randomly selected, yielding 2–5 total turns per image. Instructions are generated by Gemini-2.5-Pro to be single-context and to use referential language linking back to prior edits.

  • Positioning against prior datasets. Compared in Table 3 with GIER, MagicBrush (10K human-annotated triplets), HQ-Edit, Echo-4o-Image (approximately 180K synthetic examples), UltraEdit, OmniEdit, GPT-Image-Edit-1.5M (1.5M regenerated triplets), and UniVG, the authors state that Pico-Banana-400K emphasizes quality-controlled, instruction-faithful edits and fine-grained category coverage rather than sheer scale, and uniquely includes a 56K preference triplet subset and a diverse human-centric subset.

Methodology in Plain English

The pipeline has five practical stages.

  1. Start from real photos. Images are sampled from OpenImages, selected so that humans, objects, and textual scenes are all covered, and with category-specific filtering so human-centric and text-related edits are only attempted on appropriate images.

  2. Define what can be edited. The authors write a taxonomy of 35 edit types grouped into 8 categories: Pixel & Photometric, Object-Level Semantic, Scene Composition & Multi-Subject, Stylistic, Text & Symbol, Human-Centric, Scale, and Spatial/Layout. Each image-instruction pair gets exactly one primary edit type. Before committing, they empirically tested Nano-Banana and dropped operations that could not be rendered consistently at high quality.

  3. Write two kinds of instructions. Gemini-2.5-Flash writes a long, detailed, image-aware editing instruction for each image. Separately, human annotators wrote instructions for a subset, and those human examples were used as in-context demonstrations so Qwen2.5-7B-Instruct could rewrite instructions into short, user-style commands. Users can pick whichever variant suits them.

  4. Generate the edit, then judge it. Nano-Banana produces the edited image. Gemini-2.5-Pro then acts as an automatic judge, mimicking professional human evaluation with a weighted four-part rubric (instruction compliance, seamlessness, preservation balance, technical quality) and returning a score from 0.0 to 1.0 that is compared against the approximately 0.7 threshold. This replaced human annotators for scale, and failed attempts were retried automatically.

  5. Build the extended subsets. For multi-turn data, additional edit types are chosen per image and Gemini-2.5-Pro writes context-linked instructions, each applied to the current working image and evaluated with the same tooling. For preference data, successful and failed edits of the same instruction are paired into triplets (original image, instruction, success vs. failure).

Why This Matters

Impact on research. The paper argues that open progress in text-guided editing has been constrained by a lack of large-scale, high-quality, openly accessible datasets built from real images. By releasing both the data and the construction recipe—generation model, judge rubric, thresholds, retry policy, and taxonomy—the authors provide a reproducible template for dataset distillation, plus a single resource supporting supervised fine-tuning, preference alignment, reward modeling, multi-turn planning research, and instruction-rewriting or summarization study.

Real-world applications.

  • Photo editing and creative tools that let users describe an edit in natural language instead of manipulating layers or masks.
  • E-commerce and marketing pipelines that need consistent restyling or background/scene replacement across large image catalogs.
  • Accessibility and localization workflows, such as translating or restyling text in signs and posters, and generating concise user-style prompts from detailed ones.
  • Assistive and conversational editing interfaces where a user iterates on an image across several turns, relying on the system to keep track of prior edits.

Industry relevance. The dataset targets commercial demand for controllable image editing while highlighting where current systems still fail: precise spatial manipulation, outpainting, typography, and identity-preserving human stylization. The paper positions these as concrete engineering problems, suggesting directions such as region-referential prompting, attention steering, geometry-aware training objectives, OCR-informed losses, and identity-preserving constraints. The roughly 100K USD production cost also signals that datasets of this kind are substantial infrastructure investments, and the Apple affiliation plus public code release indicate a commercial interest in shared benchmarks.

Future Directions

  • Model benchmarking and training studies. The conclusion explicitly lists future work on benchmarking and training models with Pico-Banana-400K, and on examining how the dataset affects controllability and visual fidelity. No benchmark results are reported in this paper.

  • Fixing the hard edit types. Relocate object (0.5923), change font/style (0.5759), caricature (0.5884), and outpainting (0.6634) point to open problems in spatial conditioning, typography fidelity, and boundary continuity.

  • Identity preservation in human stylization. Pixar/Disney-like 3D (0.6463) and caricature (0.5884) show identity drift and shading artifacts under large shape exaggeration, which the authors flag as needing identity-preserving constraints.

  • Open questions the dataset enables. How much multi-turn supervision improves iterative refinement and coreference resolution; whether preference pairs from failed edits actually improve robustness under DPO or reward modeling; and how much the choice between long, detailed instructions and short, human-style instructions changes downstream model behavior.

  • Underexplored subset relationships. The paper notes that the multi-turn subset is built by expanding a sample of single-turn data, leaving open how far the multi-turn data extends beyond that sampled seed. The paper does not report whether some edit types could be added back with better models or conditioning.

Target Audience

Researchers and engineers working on text-guided image editing, multimodal generation, and dataset curation—particularly those training or fine-tuning editing models, building preference-alignment or reward-model pipelines, or studying multi-turn and planning-based editing. It is also useful for practitioners who want a shareable, license-clear training corpus, and for students who want a clear case study in automated dataset construction using a generator model, an LLM judge, and a taxonomy-driven pipeline.

Authors’ abstract

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset is constructed by leveraging Nano-Banana to generate diverse edit pairs from real photographs in the OpenImages collection. What distinguishes Pico-Banana-400K from previous synthetic datasets is our systematic approach to quality and diversity. We employ a fine-grained image editing taxonomy to ensure comprehensive coverage of edit types while maintaining precise content preservation and instruction faithfulness through MLLM-based quality scoring and careful curation. Beyond single turn editing, Pico-Banana-400K enables research into complex editing scenarios. The dataset includes three specialized subsets: (1) a 72K-example multi-turn collection for studying sequential editing, reasoning, and planning across consecutive modifications; (2) a 56K-example preference subset for alignment research and reward model training; and (3) paired long-short editing instructions for developing instruction rewriting and summarization capabilities. By providing this large-scale, high-quality, and task-rich resource, Pico-Banana-400K establishes a robust foundation for training and benchmarking the next generation of text-guided image editing models.

Read the original paper