Research
GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
Overview Research area: Human-in-the-loop optimization for generative image models — specifically, merging multiple fine-tuned diffusion "adapters" (such as LoRAs) into a single combined model by choo

- arXiv
- 2601.18585
- Published
- 2026-01-26
- Authors
- Chenxi Liu, Selena Ling, Alec Jacobson
AI summary
Overview
Research area: Human-in-the-loop optimization for generative image models — specifically, merging multiple fine-tuned diffusion "adapters" (such as LoRAs) into a single combined model by choosing per-adapter merging coefficients, using Preferential Bayesian Optimization (PBO).
Technical level: Advanced. The paper assumes familiarity with diffusion model customization, LoRA-style parameter-efficient fine-tuning, Gaussian processes, acquisition functions, and Bayesian optimization terminology.
Scope: The paper proposes GimmBO, a two-stage Bayesian optimization backend plus user interface that lets users interactively search a 20–30 dimensional space of adapter merging coefficients using only preference feedback (ranking images), rather than manually tuning sliders.
What This Paper Is About
Fine-tuning produces large libraries of community-shared adapters, and adapters built on the same base model can be merged linearly by choosing a coefficient for each adapter. Users currently explore this design space by trial and error with sliders, which the paper states scales poorly and degrades steeply beyond 5–10 sliders (Dang et al., 2022), making coefficient selection hard even when the candidate set is limited to 20–30 adapters. GimmBO instead treats merging as a continuous design problem and searches it using human preference feedback, with a two-stage optimizer that exploits two properties observed in real usage: effective merges are sparse, and merging coefficient sums are small.
Key Contributions
- A human-in-the-loop adapter merging framework (GimmBO) that casts merging coefficient selection as Preferential Bayesian Optimization, where users rank a small number of generated images per iteration rather than assigning numeric scores.
- A two-stage BO backend motivated by real-world usage statistics. It first searches a B-capped simplex (with B=2 by default) to encourage sparse, meaningful combinations, then fixes the sparsity pattern and optimizes over a smaller unconstrained hypercube for precision (T₁=10 initial-stage iterations, T₂=9 final-stage iterations).
- A constrained initialization procedure — a modified Dirichlet / "stick-breaking" process with random coordinate ordering and thresholding — to sample feasible coefficient vectors from the capped simplex, avoiding intractably slow rejection sampling as dimensionality grows.
- Extensive evaluation with simulated users (20D primary setting, plus 30D and 40D stress tests) and a 12-participant user study, plus extensions such as content control via SDEdit and integration with community model retrieval (Stylus, Luo et al., 2024).
Main Findings
-
Real-world usage is sparse. Of 20,780 Civitai images with sufficient metadata, 68% were created using at least one adapter, and of those, 69% were created by merging between two to nine adapters. Out of 3,533 users, the top 1% contributed 33.3% of works and the top 5% contributed 57.8%; the paper reports statistics remained consistent after removing those top users, indicating no significant bias.
-
Merging coefficients cluster at small magnitudes. The sum of merging coefficients falls off quickly, with most sums falling between 1.0 and 3.0. The paper notes that values within or near [0, 1] behave well and that the feasible set is fuzzy and adapter dependent, motivating a coarse proxy geometry.
-
The capped simplex is tiny in high dimensions. Vol(10; 2) ≈ 0.03% and Vol(30; 2) < 10⁻²¹ %, supporting the use of a B-capped simplex as a conservative early search space.
-
GimmBO outperforms line-search and prior BO baselines in simulation. In the 20D matching task, batch BO-based methods substantially outperformed direction-descent and coordinate-descent methods. GimmBO showed rapid Stage 1 gains from the capped simplex constraint, with further improvement at the onset of Stage 2, and consistently achieved the highest sparsity recovery (F1), with many cases reaching F1 = 1.
-
Success rates decline with more active adapters. As the ground-truth number of active adapters z increases, success rates decline for all methods, though GimmBO remains highest. The two-stage strategy may miss the global optimum (F1 ≠ 1) if Stage 1 fails to include all true active adapters.
-
Graceful degradation at higher dimensionality. In simulated stress tests at 30D and 40D, GimmBO degrades gracefully and consistently outperforms Sequential Gallery (Koyama et al., 2020).
-
Batch size and sparsity settings matter. At Step 1 and Step 10 in a 20D matching task with q=8, B=2 performed best compared with B=1 and B=n. Ablations over the simplex constraint B ∈ {1, 2, n}, top-k selection (k=1 vs. 5), and inclusion of two past samples showed the default configuration (B=2, top-5, +past) performed best; even with a matched interaction budget (B=2, top-1, −past), GimmBO substantially outperformed Sequential Gallery.
-
User study favors GimmBO. With 12 participants (9 male, 3 female; 8 aged 25–34 and 4 aged 18–24) over 20 iterations across three interfaces (slider, gallery, top-k ranking), all methods showed a statistically significant increase in similarity by the end of the session (gallery p < 0.05, slider p < 0.05, GimmBO p < 10⁻³), with GimmBO achieving the highest performance (> 0.9 on average) and higher success rates and F1, while baselines activated redundant adapters. Performance typically saturates by ~15 iterations.
-
Interaction time differs substantially. Mean interaction time per step was 10.6 s (gallery), 34.7 s (slider), and 50.5 s (ours), excluding image generation latency of ~10–15 s per batch. Under time-matched budgets (approximately 14 steps for slider and 4 steps for gallery), GimmBO outperformed slider with statistical significance (p < 0.05) and achieved higher median similarity than gallery, though that difference was not statistically significant (p = 0.397). With inference time included (15 s per iteration), GimmBO outperformed both baselines with statistical significance (p < 0.05); at 60 s inference, median similarity was 0.956 vs. 0.898 for slider and 0.926 vs. 0.854 for gallery.
-
Content control adds little style cost. Using SDEdit (Meng et al., 2021) at control strength 0.2, style similarity was 0.77 (where 1 is a perfect match) in the default setting.
-
Integration with retrieval produces diverse results. Given a text prompt, 20–25 community-shared adapters gathered via coarse text-based retrieval of Stylus (Luo et al., 2024) were compared across trivial averaging, Stylus refinement, and GimmBO: averaging often introduced artifacts, Stylus-picked configurations yielded a single fixed result, while GimmBO produced multiple, diverse, prompt-aligned outcomes per prompt through user-guided exploration (examples generated by SD1.5).
-
Prompt-based alternatives struggle on underrepresented styles. Given 30 reference images by an artist with few publicly available works, fine-tuning an SDXL LoRA closely matched the target style, whereas ChatGPT image creation, auto-captioning followed by SDXL generation, and PromptMagician (Feng et al., 2023) achieved only partial similarity.
Methodology in Plain English
The problem is framed as maximizing an unknown user utility function f over generated images g(α), where α is the vector of merging coefficients. Because the user's taste is a black box and every evaluation costs a model inference plus human time, the authors use Bayesian optimization, which builds a cheap surrogate model of the user's preferences and uses it to decide which images to show next.
Instead of asking users for numeric ratings, GimmBO asks them to rank a small batch of images. Ranking N images and marking the top k yields up to k(k−1)/2 + k(N − k) pairwise inequalities, which are converted into latent function values using the preference learning approach of Chu and Ghahramani (2005), and then fit with a Gaussian process using a sparse axis-aligned subspace (SAAS) prior (Eriksson and Jankowiak, 2021). New samples are chosen with an upper confidence bound (UCB) acquisition function (Srinivas et al., 2009), extended to a batch version that maximizes the expected best single acquisition value (Wilson et al., 2017). Sampling 8 coefficient vectors per iteration.
Two domain observations shape the search space. First, people merge only a few adapters, so the search begins in a B-capped simplex (coefficients non-negative and summing to at most B, with B=2 by default) rather than the full [0, 1]ⁿ hypercube. Second, the sum of coefficients must stay small to avoid degraded generations. Samples are drawn from this simplex using a modified Dirichlet / stick-breaking process with random coordinate ordering, then thresholded to increase sparsity. After T₁=10 iterations, the sparsity pattern of the current best coefficient vector is fixed, and the remaining T₂=9 iterations optimize only the surviving coefficients in an unconstrained space.
For content preservation, SDEdit is applied with a reference image generated by the base model, so users judge style rather than being distracted by content drift. The frontend presents N images pre-sorted by surrogate estimate and asks users only to identify and order the top k, with the selected image shown at full resolution. The backend is implemented in Python using BoTorch; the paper directs readers to the supplemental material for detailed hyperparameter values.
Evaluation uses a "matching task" standard in human-in-the-loop BO: a target coefficient vector produces a ground-truth image visible only to the simulated user (or withheld from the participant in the human study), and the method must recover it. The default protocol allows 5 initial inferences followed by 20 optimization iterations with 8 inferences each. Experiments use collections of n=20 adapters assembled from popular Civitai LoRAs, m=30 prompts (five runs per prompt with different random seeds) of the form "a drawing of <noun>" or "a portrait of <noun>" for person nouns, with nouns drawn from CIFAR-10 and CIFAR-100 super-classes. Result quality is measured with DreamSim (Fu et al., 2023), normalized to [0, 1], where 0.9 and 0.95 correspond empirically to rough and near-exact matching; adapter recovery is measured with F1.
Why This Matters
Impact on research. The paper argues that existing BO-based interactive design methods were built for 3–12D spaces and are ill suited to 20–30D adapter merging — Sequential Slider (Koyama et al., 2017) relies on single-direction updates that do not scale, and Sequential Gallery (Koyama et al., 2020) assumes interior solutions and ignores poor-quality regions, producing clamped or degraded samples. GimmBO shows that injecting domain-specific structure (sparsity, bounded coefficient sums) into the BO backend improves sampling efficiency and convergence, a pattern likely transferable to other high-dimensional, preference-driven design problems.
Real-world applications:
- Interactive style exploration on community image-generation platforms, where users blend many shared adapters.
- Personalized content creation workflows that need to combine a specific subject adapter with several style adapters while preserving content.
- End-to-end pipelines that pair automatic adapter retrieval (e.g., Stylus) with guided exploration of the retrieved 20–30 candidates.
- Design and creative tooling in general, since the optimization backend places no requirements on the image generator g and could support other parameterized generators.
Industry relevance. Platforms such as Civitai already see heavy multi-adapter usage (69% of adapter-based generations in the December 2025 statistics cited), but offer only manual slider control. GimmBO's argument that its advantage grows when inference is slow (with 60 s inference it reported 0.956 vs. 0.898 for slider and 0.926 vs. 0.854 for gallery) speaks directly to users running limited local GPUs or queued online jobs, where inference latency dominates interaction time.
Future Directions
- Failing Stage 1 is a hard ceiling. The two-stage strategy can miss the global optimum (F1 ≠ 1) when the initial stage never includes all true active adapters, so recovering from a wrong sparsity pattern is an open problem.
- Integrating interface design work. The authors note that cooperative design-process methods (Koyama and Goto, 2022; Mo et al., 2024; Niwa et al., 2025) emphasize interface design and are complementary to their optimization backend, calling integration an interesting future direction.
- Beyond style merging. The paper demonstrates extensions to content/subject merging and downstream applications, but the quantitative evaluation focuses on style adapters, described as less structured and under-explored relative to style-content composition.
- Complexity of human preferences. The authors note that the human-preference objective is unlikely to be additive, which limits the applicability of embedding, trust-region, and other high-dimensional BO mechanisms without adaptation — leaving room for better-suited solvers.
Target Audience
Researchers and practitioners in computer vision, graphics, and human-computer interaction working on generative image customization, model merging, and human-in-the-loop design optimization. It is also relevant to engineers building creative tools on top of diffusion model adapter libraries, and to methodologists interested in preferential Bayesian optimization in high-dimensional, sparsely structured search spaces. Readers without background in Gaussian processes, acquisition functions, or LoRA-style fine-tuning will find the methodology sections demanding.
Authors’ abstract
Fine-tuning-based adaptation is widely used to customize diffusion-based image generation, leading to large collections of community-created adapters that capture diverse subjects and styles. Adapters derived from the same base model can be merged with weights, enabling the synthesis of new visual results within a vast and continuous design space. To explore this space, current workflows rely on manual slider-based tuning, an approach that scales poorly and makes weight selection difficult, even when the candidate set is limited to 20-30 adapters. We propose GimmBO to support interactive exploration of adapter merging for image generation through Preferential Bayesian Optimization (PBO). Motivated by observations from real-world usage, including sparsity and constrained weight ranges, we introduce a two-stage BO backend that improves sampling efficiency and convergence in high-dimensional spaces. We evaluate our approach with simulated users and a user study, demonstrating improved convergence, high success rates, and consistent gains over BO and line-search baselines, and further show the flexibility of the framework through several extensions.