Research
VibeEdit: Image Editing with Canvas Instructions
Overview Research area: Computer vision, specifically generative image editing with diffusion models and instruction-following editors. Technical level: Advanced. The paper assumes familiarity with di
- arXiv
- 2610.12229
- Published
- 2026-10-08
- Authors
- Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu
AI summary
Overview
- Research area: Computer vision, specifically generative image editing with diffusion models and instruction-following editors.
- Technical level: Advanced. The paper assumes familiarity with diffusion transformers, vision–language encoders, LoRA adapters, and reinforcement learning post-training for generative models.
- Scope: The paper proposes a new on-image editing interface called "canvas instructions," builds a training corpus and benchmark for it, and trains VibeEdit to perform five edit types from those instructions without a separate text prompt.
What This Paper Is About
Text-guided image editors require users to describe both what to change and which object or region to change. Describing the change is usually easy, but pointing at the right object is hard when several objects look alike. VibeEdit instead lets users draw directly on the image — circles, scribbles, arrows, plus optional short handwritten notes — and treats those marks as the complete editing instruction.
Key Contributions
-
A new editing interface. Spatial marks (circles, scribbles, drag gestures such as arrows) plus optional short text notes placed on the image form a "canvas instruction" that specifies both the edit target and the desired change across five operations: addition, removal, replacement, attribute modification, and movement. No separate text prompt is needed.
-
A large training corpus. The authors construct 1.55 million source–target edit pairs, each with object masks and structured edit descriptions. Canvas annotations are rendered online during training by sampling stroke shapes, widths, and typefaces (including handwriting and print styles), so the same underlying edit can be presented with different annotations.
-
A model architecture and training recipe. VibeEdit adapts Qwen-Image-Edit with layer-decoupled conditioning (the vision–language encoder reads the annotated image while the VAE separately encodes the clean source image and the rasterized annotation layer), region-weighted supervised fine-tuning, and rubric-guided reinforcement learning.
-
An independent benchmark. A human-curated benchmark of 419 cases, built independently of the training corpus, emphasizing scenes with multiple similar objects, with human-written external text instructions (21.3 words on average) prepared so text-instructed baselines can be compared fairly.
Main Findings
-
VibeEdit leads on both reported metrics. VibeEdit achieves an overall VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB on the 419-case benchmark, versus 67.4 and 24.0 dB for FireRed-Image-Edit-1.0, the highest-scoring text-instructed baseline in the evaluation.
-
It wins on every individual edit type. VibeEdit scores 80.3 / 34.7 (add), 85.4 / 35.8 (remove), 77.4 / 31.2 (attribute), 79.9 / 34.9 (replace), and 76.4 / 27.1 (move), reporting the highest rubric score on all five types.
-
The base model is already competitive with the best text-instructed baseline. VibeEdit-base reaches an overall VLM rubric score of 67.8 and PSNR of 27.0 dB, against 67.4 / 24.0 dB for FireRed with external text instructions.
-
Text instructions help all baselines. All six public baselines achieve higher overall rubric scores with external text instructions than with the canvas-annotated image alone. For example, Qwen-Image-Edit-2511 goes from 19.8 / 16.3 in the visual-only setting to 62.5 / 19.4 with text instructions, and JoyAI-Image-Edit from 46.4 / 24.5 to 65.4 / 23.3.
-
Layer-decoupled conditioning matters. With the vision–language input fixed to the annotated image, encoding only the source image as VAE input yields a 24.7 rubric score and 21.3 dB; encoding the annotated image yields 58.6 / 24.2; separately encoding source and rasterized annotations yields 67.8 / 27.0.
-
Region weighting needs dilation to help. Region weighting without mask dilation lowers the rubric score from 66.2 to 64.2, while weighting combined with dilation raises it to 67.8.
-
High-variance RL selection is best. Selecting RL conditions by rollout-reward variation gives 79.9, ahead of random selection at 79.3 and low-variance selection at 76.6.
-
All three reward groups are needed. Edit-success rewards alone give 69.2 / 28.6 dB; adding outside preservation gives 74.9 / 29.4 dB; adding local edit quality instead gives 72.4 / 28.0 dB; combining all three gives 79.9 / 32.8 dB.
-
Joint training transfers across tasks. Training on all five tasks yields an overall 67.8. Single-task models reach 45.2 (add only), 47.8 (remove only), 45.9 (attribute only), 56.1 (replace only), and 38.0 (move only). Removal-only and replacement-only models outperform the joint model on their own tasks (81.4 on removal and 77.6 on replacement, respectively), but joint training improves over single-task models on addition, attribute modification, and movement.
-
Stroke color is not a strong dependency. Training uses red strokes and green movement destination circles. Relative to red, blue gives 79.4 / 32.6, yellow 79.3 / 32.6, and per-case random sampling from red, blue, and yellow gives 79.5 / 32.6 — changes of at most 0.6 rubric points and 0.2 dB.
-
Font diversity improves human preference. Across 200 pairwise human comparisons of edits guided by user-written canvas instructions, the model trained with multiple fonts scores an Elo of 1617.1 versus 1382.9 for the single-font model.
-
Removal works best with a scribble plus a text cue. On the removal subset: X mark 58.4 / 27.0, X mark with text 78.9 / 30.8, scribble 55.4 / 27.6, scribble with text 81.4 / 36.1. The authors adopt a circled scribble with "remove it" as the default.
-
The automatic judge tracks human labels. On 100 randomly sampled VibeEdit outputs and 947 matched binary rubric questions, pooled pass rates are 76.6% for humans and 77.0% for GPT-5.6-sol, with 93.0% accuracy, 95.5% F1, and Cohen's κ = 0.805 at question level.
-
Attention plus modulation LoRA is the right scope. Attention-only adapters (377.5M trainable parameters) give 64.1 / 26.6; attention plus modulation (707.8M) gives 67.8 / 27.0; adding the MLP (943.7M) drops to 60.3 / 22.7.
Methodology in Plain English
The authors separate the underlying edit from the way it is drawn. Offline, they build a corpus of source–target image pairs. Florence-2 proposes candidate objects, GPT-5.6-sol picks one that is recognizable, visible, and separable and writes an edit description, and SAM 3 produces the object mask. Operation-specific pipelines then generate the edited target: ObjectClear inpaints a masked region to produce removal pairs (reversing them gives addition pairs); Qwen-Image-Edit applies attribute changes; FLUX.1 Fill inserts replacement objects; and Qwen-Image-Edit relocates objects for movement, with SAM 3 re-locating the destination mask. GPT-5.6-sol screens the pairs for quality. The corpus has 1,554,062 stored image pairs.
During training, canvas instructions are rendered on the fly: circles and short notes for addition, attribute modification, and replacement; a circled scribble with a deletion note for removal; two circles joined by a directed arrow for movement.
The model starts from Qwen-Image-Edit. A frozen Qwen2.5-VL encoder reads the annotated image for semantic interpretation. Separately, the VAE encodes the clean source image and, on its own neutral gray background, the rasterized annotations. The DiT therefore receives noisy target tokens, source tokens, canvas tokens, and semantic tokens. Rank-128 LoRA adapters are trained on the DiT attention projections and modulation layers while the VAE, vision–language encoder, and DiT weights stay frozen.
Supervised fine-tuning upweights the loss inside a dilated mask covering the edited object and the rendered annotations (λ = 1.5, dilation of 50 pixels before downsampling), so the model learns both to complete the edit and to erase the marks. The mask only weights the loss; it is not given to the model at inference.
Reinforcement learning uses DiffusionNFT with a rubric reward. For each RL condition, the old policy samples 12 rollouts and a fixed GPT-5.4-mini judge answers binary questions in three groups — edit success, outside preservation, and local edit quality. A penalty of γ_p = 0.3 is applied when outside-region PSNR falls below τ_p = 25 dB. Rewards are centered within each condition and normalized by the edit-type reward standard deviation. The RL set of 3,520 conditions (704 per edit type) is chosen for high reward variation among rollouts.
Training details: supervised learning uses aspect-ratio buckets of roughly 1024² pixels, global batch size 128, and a constant learning rate of 5×10⁻⁵ after 500 warmup steps. RL starts from the 12,000-step checkpoint at 512² resolution and is evaluated at 220 steps. Both stages run on 32 NVIDIA A100 GPUs. At inference, base models use 30 denoising steps and CFG scale 4; RL models use 16 steps and CFG scale 1.
Why This Matters
The paper reframes where the difficulty in image editing actually sits. Modern models are good at understanding what change is requested; they are weaker at knowing which object the user means. By moving the reference step from prose to a mark on the image, VibeEdit removes the need for the contorted spatial descriptions ("the leftmost of the two identical chairs in the back") that text prompting forces.
- Research impact: It introduces a distinct interface-plus-data-plus-training formulation, an independently built benchmark of 419 cases emphasizing similar-object target selection, and evidence that layer-decoupled conditioning and decomposed rubric rewards transfer to localized editing.
- Consumer photo editing: Direct annotation on a photo suits mobile and touch interfaces, where drawing a circle is faster than typing a sentence.
- E-commerce and product photography: Removing, replacing, or restyling objects while preserving the rest of the frame maps onto catalog and listing workflows.
- Design and marketing: Several similar-looking assets in one scene can be edited by pointing rather than describing position.
- Industry relevance: The recipe is built on top of an existing open base editor (Qwen-Image-Edit) with LoRA adapters, so the approach is portable to other hosted editors and shows a path to spatial grounding without replacing the generative backbone.
Future Directions
- Scaling the corpus and the benchmark. The corpus has 1,554,062 pairs and the benchmark has 419 cases; whether the interface holds up on harder or larger-scale evaluation is not established here.
- Reducing dependence on synthesized supervision. Training targets are generated by other models, and the authors note that filtering may not eliminate all errors — they rely on RL to mitigate residual training noise rather than claiming clean data.
- Robustness beyond stroke color and font. The paper tests stroke color variation and font diversity, but other factors such as mark thickness conventions, canvas clutter, or unusual user handwriting are not explored.
- Instruction grammar and interaction vocabulary. The paper reports four removal designs and defines circle, scribble, and arrow conventions; how far this grammar extends to more complex, multi-object, or compositional edits remains an open question.
Target Audience
Researchers and practitioners working on diffusion-based image editing, multimodal interfaces, and reinforcement learning for generative models. It is most useful to readers who already understand diffusion transformers, VAE latent conditioning, and RL post-training, and who want to see how spatial grounding can be folded into an editor's conditioning rather than its text prompt. Product and design teams building creative tools will find the interface framing and the ablation results on instruction design most directly actionable.
Authors’ abstract
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.