Research
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
Overview Research area: Multimodal image generation, specifically agentic "harness" design for multi-reference image generation (combining several reference images into a new image), combined with aut

- arXiv
- 2609.35530
- Published
- 2026-09-28
- Authors
- Yuta Oshima, Ku Onoda, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
AI summary
Overview
Research area: Multimodal image generation, specifically agentic "harness" design for multi-reference image generation (combining several reference images into a new image), combined with automatic optimization of agent code.
Technical level: Advanced. The paper assumes familiarity with image generation models, vision-language models (VLMs), agentic refinement loops, beam search, and benchmark evaluation protocols.
One-sentence scope: The paper introduces AutoRef, a search procedure in which a coding agent iteratively rewrites the executable "harness" that orchestrates a frozen image generator and a frozen reasoning model, and reports AutoRef-Harness, the harness it discovers for four-reference MultiBanana tasks.
What This Paper Is About
Multi-reference image generation models often omit or duplicate subjects from the references, or produce images where multiple subjects look copied and pasted rather than belonging to one coherent scene. Wrapping such a model in an "agent" requires writing a harness — executable code that decides how references are interpreted, how prompts are built, how candidates are generated and evaluated, and how the final image is chosen — but many parts of that harness could be changed, the effect of each change is hard to predict, and human-written harnesses vary widely in performance. AutoRef addresses this by automatically optimizing the harness code while keeping both the image generator and the reasoning model frozen.
Key Contributions
-
AutoRef, a harness optimization method for multi-reference image generation. It keeps all model parameters frozen and changes only the executable program around them, targeting perceptual image-generation tasks where rewards are noisy and non-verifiable rather than discrete or executable-test-verifiable.
-
Task separation between proposal and selection. Search tasks are split into disjoint
D_trainandD_valsets: everything onD_train(scores, evaluator rationales, execution trajectories, and generated images, including from rejected candidates) enters the search history that the proposer may inspect, whileD_valonly ranks candidates and is never exposed to the proposer. -
Iterative beam search over harnesses. Rather than committing to a single lineage, AutoRef generates
Kcandidate harnesses per iteration and keeps the topBon the validation tasks as parents for the next iteration (in the experiments,B = 2andK = 4). -
AutoRef-Harness, the discovered harness, plus transfer evidence. The paper releases the harness and code at https://github.com/KuOnoda/AutoRef, and shows the same harness improves results without re-optimization when the generator, number of references, benchmark, evaluator, or reasoning model changes.
Main Findings
-
Base generator improved substantially on the four-reference MultiBanana held-out test split. AutoRef-Harness raises the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10 on the average of the reported metrics), which matches or exceeds proprietary models including GPT-Image-1.5 (7.07), Nano Banana Pro (7.20), and Seedream 4.5 (7.03), and outperforms all evaluated open models (OmniGen2 3.75, DreamOmni2 3.83, BAGEL 3.15, Qwen-Image-Edit-2511 4.44).
-
The harness uses three image generations per task. The "Gen." column is a measured mean over the split, and AutoRef-Harness draws 3 images per task.
-
Transfer across generators without re-optimization. Applied unchanged, the harness improves FLUX.2 [klein] 9B from 5.78 to 7.10 and Qwen-Image-Edit-2511 from 4.44 to 5.69 on the four-reference test split. It also improves Qwen-Image-Edit-2511 after fine-tuning for multi-reference generation with DyRef.
-
Transfer across reference counts. On unseen three-reference MultiBanana tasks, FLUX.2 [klein] 4B goes from 6.94 to 7.76 with the harness; FLUX.2 [klein] 9B from 6.83 to 7.82; Qwen-Image-Edit-2511 from 5.06 to 6.95. On unseen five-reference tasks, FLUX.2 [klein] 4B goes from 5.19 to 6.36 and FLUX.2 [klein] 9B from 5.61 to 6.47. The exception is Qwen-Image-Edit-2511 at five references, where the generator itself fails and the harness cannot compensate (1.89 without the harness vs. 1.87 with it).
-
Transfer across benchmark and evaluator. On OmniContext, never used during the search and scored by its official GPT-4.1 evaluator rather than the MultiBanana Qwen3-VL-8B-Instruct evaluator, the harness improves FLUX.2 [klein] 4B from 8.30 to 8.85, FLUX.2 [klein] 9B from 8.59 to 8.97, and Qwen-Image-Edit-2511 from 8.45 to 8.70.
-
Transfer across reasoning models. Replacing GPT-5.5 with the open-weight Qwen3-VL-32B while leaving the harness structure unchanged improves FLUX.2 [klein] 4B from 5.72 to 6.90 on the four-reference test split (7.37 with GPT-5.5).
-
Outperforming human-written harnesses. AutoRef-Harness achieves the highest performance among Best-of-3, GEMS, IPR, and Idea2Img on the four-reference held-out test split. These use 3, 2.7, 3, and 9 image generations per task respectively, so Idea2Img still underperforms despite using three times AutoRef-Harness's generation budget.
-
Outperforming alternative search procedures. Under the same generator, reasoning model, and evaluator, AutoRef yields the strongest final harness, followed by Greedy Search and then Meta-Harness. The final harnesses of Meta-Harness and Greedy Search also draw more images per task (4.2 and 5 vs. 3), and AutoRef's second-best harness outperforms both methods' final harnesses.
-
Every harness component contributes (ablation on the held-out test split). Full harness 7.37; removing structurally diverse drafts 7.01 (−0.36); removing complaint-directed revision 6.99 (−0.38); removing failure-aware selection 6.93 (−0.44); grounding only 6.91 (−0.46); selection only 6.15 (−1.22); generator only 5.72 (−1.65). Grounding only (6.91, one image) approaches IPR (7.02, three images), and selection only (6.15) is close to Best-of-3 (6.01), yet the same selector adds 0.44 inside the full harness.
-
Human raters prefer the harness output. Four raters compared FLUX.2 [klein] 4B + AutoRef-Harness with four baselines on 50 tasks each from the four-reference held-out test split. The harness wins 70% against the base FLUX.2 [klein] 4B (17% loss), beats FLUX.2 [klein] 9B (64% vs. 24%) and Seedream 4.5 (63% vs. 27%), and is competitive with Nano Banana Pro (46% vs. 40%).
Methodology in Plain English
The setup starts from two frozen models: an image generator and a reasoning (vision-language) model. The "harness" is ordinary code that decides how those two models are called — how references are interpreted, how prompts are written, how many images to generate, how outputs are inspected, and which one is returned.
AutoRef improves that code with a coding agent as the proposer. At each iteration the proposer reads the current set of harnesses plus a history of previous attempts (implementations, scores, execution trajectories, and the generated images themselves) and writes new harness candidates. Crucially, the tasks are split in two: a training set whose rich feedback (scores, evaluator rationales, trajectories, images) is visible to the proposer and drives its revisions, and a separate validation set used only to rank candidates. The proposer is told which candidates survived but never their validation scores, so the selection signal cannot be gamed directly.
Instead of keeping only the single best harness at each step, AutoRef keeps a beam of the top B harnesses on the validation tasks and uses them as parents for the next round, generating K candidates per iteration. In the experiments, the search ran five iterations, was initialized with the plain FLUX.2 [klein] 4B generator and the GEMS harness, and used Claude Fable 5.1 through the Claude Code CLI as the proposer. Optimization was done on four-reference MultiBanana tasks (48 training, 48 validation, 133 held-out test tasks) with Qwen3-VL-8B-Instruct as the evaluator and GPT-5.5 as the reasoning model.
The resulting harness draws three images: it generates drafts A and B from two differently structured prompts, keeps the better one, generates draft C from explicit complaints about the winner, and returns the better of the winner and C.
Why This Matters
Impact on research. The paper reframes "how frozen models are used" as an object worth optimizing in its own right, and shows that this is distinct from prompt tuning within a fixed program or from updating model weights. It also provides evidence that train/validation separation and a multi-parent beam matter specifically for perceptual, non-verifiable tasks such as image generation, where scalar rewards carry little diagnostic information. The reported gap between AutoRef, Greedy Search, and Meta-Harness isolates the search algorithm as the only difference under matched budget, generator, reasoning model, and evaluator.
Real-world applications (drawn from the paper's stated motivations):
- Advertising, where users specify people, objects, and backgrounds using separate images.
- Virtual try-on, where garments and identities come from different references.
- General content creation requiring controllable composition of multiple subjects and a scene.
- Converting an instruction plus multiple references into a single coherent image for downstream creative workflows.
Industry relevance. The headline result is that a small open-weight generator (FLUX.2 [klein] 4B) combined with an automatically discovered harness can match or exceed proprietary systems such as Nano Banana Pro and GPT-Image-1.5 on this benchmark, using only inference-time code changes and three generated images per task. That suggests a route to quality gains without retraining or fine-tuning models, and it gives practitioners a reusable harness that transfers to different generators, reference counts, benchmarks, evaluators, and reasoning models without re-optimization.
Future Directions
- Combine the harness with stronger or richer evaluators. The paper notes that the evaluator and proposer are replaceable, so AutoRef may benefit from stronger VLMs and coding agents, and suggests richer evaluators such as combining a VLM with segmentation models.
- Address the ceiling imposed by the generator. Because the harness changes only how a frozen generator is used, it cannot exceed what that generator can produce; the Qwen-Image-Edit-2511 five-reference case is given as an example where better selection does not help.
- Test harness optimization under other evaluation regimes. The search maximizes a single VLM evaluator's score, so whether the same procedure holds for other evaluators, task types, or multi-criteria objectives remains open.
- Explore combinations with model-side improvements. The paper shows the harness adds to gains from fine-tuning with DyRef, raising the question of how far harness optimization and model training can be stacked.
Target Audience
Researchers and engineers working on multimodal generative agents, test-time scaling, and agentic image generation or editing; practitioners who need controllable multi-reference image composition; and anyone interested in automatic optimization of executable agent code for tasks whose rewards are perceptual rather than verifiable. Readers focused on low-level generative modeling, diffusion architecture design, or theoretical optimization will find less direct relevance.
Authors’ abstract
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.