Research
Amortized Active Generation of Pareto Sets
Overview Research area: Machine learning — multi-objective black-box optimization, generative modeling, and Bayesian optimization for discrete sequence design. Technical level: Advanced. The paper bui

- arXiv
- 2510.21052
- Published
- 2025-10-23
- Authors
- Daniel M. Steinberg, Asiri Wijesinghe, Rafael Oliveira, Piotr Koniusz, Cheng Soon Ong, Edwin V. Bonilla
AI summary
Overview
Research area: Machine learning — multi-objective black-box optimization, generative modeling, and Bayesian optimization for discrete sequence design.
Technical level: Advanced. The paper builds on variational inference, class probability estimation, and Pareto dominance theory, and assumes familiarity with ELBO-style objectives and multi-objective optimization terminology.
Scope: The paper introduces A-GPS (active generation of Pareto sets), an online framework that learns an amortized conditional generative model of the Pareto set for discrete black-box multi-objective optimization, with a-posteriori user-preference conditioning and no explicit hypervolume computation.
What This Paper Is About
Many scientific and engineering problems require optimizing several conflicting objectives over discrete objects (for example, protein sequences) where each evaluation is expensive, so the number of queries must be small. Traditional multi-objective Bayesian optimization handles this with acquisition functions such as expected hypervolume improvement (EHVI), which require complex numerical integration and scale poorly as the number of objectives grows, while scalarization approaches depend on sampling density to capture complex Pareto front geometries. A-GPS instead learns a generative model of the Pareto set directly and online, guided by a classifier that predicts non-dominance, so that users can later sample designs matching their own trade-off preferences without retraining.
Key Contributions
-
The A-GPS framework. A new formulation for online discrete black-box multi-objective optimization that recasts the problem as active learning of a generative model of the Pareto set, extending the variational search distributions (VSD) framework from single-objective to multi-objective optimization via a Pareto-set class probability estimator (CPE).
-
A theoretical link between non-dominance and hypervolume. The paper proves (Theorem 1) that for every point not in the current observable Pareto set, the hypervolume improvement (HVI) indicator equals the non-dominance indicator, and consequently (Corollary 1) that a CPE trained on non-dominance labels estimates the probability of hypervolume improvement (PHVI). This lets the method avoid explicit hypervolume computation.
-
Amortized preference conditioning. The introduction of preference direction vectors — unit vectors
uin{u ∈ R^L : ||u||_2 = 1}derived asu_n = (y_n − r)/||y_n − r||_2, whereris a reference point — plus an alignment indicatora ∈ {0,1}. Together these let a trained model sample across the Pareto front conditioned on user-specified trade-offs without retraining, in contrast to scalarization, where each new weight vector would require retraining. -
A practical training scheme. An amortized ELBO objective (A-ELBO) combining a Pareto CPE and an alignment CPE, plus an off-policy gradient estimator with importance weights that the authors report is far faster than on-policy methods such as REINFORCE for complex variational families like causal transformers.
Main Findings
-
Non-dominance classifiers implicitly estimate PHVI. Theorem 1 establishes that
1[HVI(x) > 0] = z(x)for everyxnot in the current observable Pareto set (under the assumption of zero observation noise), so a CPE trained with a proper loss on non-dominance labels predicts the probability of hypervolume improvement. The authors note a discrepancy for points already in the Pareto set, wherez(x) = 1butHVI(x) = 0, and argue that for continuous objectives this is almost surely equivalent to a held-out construction because there will be no ties. -
Preference directions generalize scalarization weights. Any convex scalarization weight
λcan be mapped to a unit vectoru = λ/||λ||_2, and eachuinduces a unique normalized weight, so conditioning on(u, a)is both more flexible and closer to the non-dominance structure than scalarizing byλ. The authors connect this to "reference vectors" used in many-objective evolutionary algorithms. -
Alignment is enforced with contrastive data. The alignment CPE is trained on aligned pairs
(a_n = 1, x_n, u_n)plus purposefully misaligned pairs formed by permutations(a_n = 0, x_n, u_{ρ_i(n)}). Two permutation schemes are used: random permutations, which cover the space of misalignment, and top-knearest neighbors by cosine distance onu_n, which improves angular precision. All experiments use 7 random permutation replicates and 2 top-2 replicates, for a total ofP = 9. -
Off-policy gradients are used instead of on-policy gradients. The paper reports that REINFORCE-style on-policy gradient estimation is very slow for complex variational distributions such as causal transformers, because low learning rates are needed to avoid exploding gradients on long sequences and new samples must be drawn at every SGD iteration. The off-policy estimator uses importance weights
w(x,u) = q_φ(x|u)/q_{φ'}(x|u), withφ'updated every 100 iterations of optimizing A-ELBO, or when the effective sample size drops below0.33S. -
Label annealing avoids overly exploitative behavior. Using the current Pareto set
S_Pareto^tdirectly as the label set can be too exploitative, so labels are defined over the union of Pareto ranks up to a thresholdτ_t(following the Pareto ranking method of the cited prior work), withτ_Tlabeling justS_Pareto^t. -
A design choice for the KL coefficient. All experiments set
β = 0.5, because the authors report that the full KL regularization (β = 1, which corresponds to exact minimization of the reverse KL objective) can hamper exploitation in later rounds on some tasks. -
Evaluation is on synthetic and protein-design tasks, measured by relative HVI. The primary performance measure is Pareto front relative HVI, using the cited implementation. The paper states that A-GPS "performs well against competing approaches on a suite of challenging synthetic and real multi-objective optimization benchmarks," and that results on synthetic benchmarks and protein design tasks demonstrate strong sample efficiency and effective preference incorporation.
-
Synthetic test functions used include negative Branin-Currin and DTLZ7. Negative Branin-Currin is described as
D = 2,L = 2, and DTLZ7 asD = 7— the provided content is truncated mid-sentence at this point. A third test function is mentioned but its identity and parameters are not present in the provided content. Numerical results for these benchmarks are not reported in the provided content. -
Three high-dimensional sequence design challenges are used. These emulate real protein engineering tasks. Their names, dataset sizes, and quantitative outcomes are not reported in the provided content.
-
Position relative to prior methods. In the comparison table, A-GPS is the only listed method marked as designed for MOO, online black-box optimization, amortized preference conditioning, non-convex Pareto fronts, discrete/mixed design spaces, a generative preference model
q_γ(u), and a modular generative observation modelq_φ(x)guided by a dominance CPE.
Methodology in Plain English
The approach works in rounds. At each round the method keeps a dataset of designs that have already been evaluated by the expensive black box, producing noisy objective vectors.
First, the observed designs are ranked by Pareto dominance: a design is labeled positive if it is not dominated by any other observed design, and the label set is annealed over rounds so that early rounds include broader non-dominated layers and the final round includes only the Pareto set. A discriminative model — the Pareto CPE — is trained to predict this label from a design and a preference direction.
Second, preference directions are computed from each observed objective vector by subtracting a reference point and normalizing to unit length. To teach the model what "aligned" means, the method builds contrastive training data: correct (design, direction) pairs are labeled aligned, and pairs where the directions have been shuffled across designs are labeled misaligned. An alignment CPE is trained on this data, so it learns to score whether a design genuinely matches a requested trade-off.
Third, a small generative model over preference directions is fit, either unconditionally in the first round to aid exploration or by maximum likelihood over the directions of the Pareto-set designs.
Fourth, a conditional generative model over designs — conditioned on the preference direction — is trained to maximize the amortized ELBO. The objective rewards generating designs that the Pareto CPE thinks are non-dominated and that the alignment CPE thinks match the requested direction, while a KL term keeps the generator close to the prior over the design space. Because backpropagating through sampling is unstable for long sequences, gradients are estimated off-policy with importance weights, reusing samples until the effective sample size degrades.
Finally, candidate designs are drawn from the generator — either using a user-specified preference direction or directions sampled from the learned direction model for broad front exploration — these are evaluated by the black box, added to the dataset, and the loop repeats. Because the generator is conditioned on the direction, a new preference can be supplied after training without any retraining.
Why This Matters
Impact on research. The paper provides a formal equivalence between non-dominance indicators and hypervolume improvement indicators, which connects two previously separate lines of work: dominance-based classification and hypervolume-based acquisition. It also removes the need for explicit hypervolume computation and scalarization in a generative multi-objective optimizer, and it extends generative "active generation" methods from single-objective to multi-objective settings while adding amortized preference conditioning that prior generative frameworks lacked.
Real-world applications (all drawn from the application domains the paper names):
- Protein engineering, where practitioners may want to jointly improve thermal stability, catalytic turnover rate, and expression yield, knowing that improving one property can degrade another.
- Small molecule synthesis with tailored pharmacokinetic properties.
- DNA construct engineering for precise gene regulation.
- Any discrete or mixed discrete-continuous design problem where each candidate must be assessed through expensive simulations or laboratory assays.
Industry relevance. The paper targets settings where evaluation is the bottleneck — computationally intensive simulations or wet-lab assays — so sample efficiency translates directly into cost and time savings. The amortized preference conditioning is particularly relevant to decision-making workflows: stakeholders often cannot state their trade-offs up front, and this method lets them explore different trade-off directions from a single trained model rather than paying for a fresh optimization run each time.
Future Directions
-
Scaling to more objectives. The paper motivates the work by the poor scaling of hypervolume-based acquisition functions with the number of objectives; how A-GPS behaves as
Lgrows is a natural open question, and the provided content does not report a scaling study. -
Handling ties and observation noise. The equivalence between non-dominance and HVI indicators carries a caveat for points already in the Pareto set (
z(x) = 1butHVI(x) = 0). The authors argue this is almost surely benign for continuous objectives and suggest that with observation noise it can be beneficial to re-sample at the same design for noise reduction — but a full treatment of noisy, tied regimes is left open. -
Improving the variational families. The paper notes that Normal distributions normalized to the unit sphere were found more numerically stable than specialized spherical distributions such as von Mises-Fisher or power spherical distributions, and that unconditional direction fitting in the initial round aids exploration. Better direction models remain a design question.
-
Broader benchmarking and comparison. The paper's evaluation covers continuous synthetic test functions, for which A-GPS was not specifically designed, plus three sequence design challenges. Extending to other discrete and mixed domains, and to a broader set of baselines than those compared, would test generality.
Target Audience
Researchers and practitioners working on multi-objective black-box optimization, Bayesian optimization, and generative approaches to design — particularly those working with expensive black-box evaluations over discrete or mixed discrete-continuous spaces such as protein and DNA sequence design. It is also relevant to readers interested in preference-conditioned generative models, amortized variational inference, or the theoretical relationship between dominance-based classification and hypervolume-based acquisition. A working knowledge of variational inference and Pareto dominance concepts is assumed; readers without that background will find the technical sections demanding, though the introductory framing is accessible.
Authors’ abstract
We introduce active generation of Pareto sets (A-GPS), a new framework for online discrete black-box multi-objective optimization (MOO). A-GPS learns a generative model of the Pareto set that supports a-posteriori conditioning on user preferences. The method employs a class probability estimator (CPE) to predict non-dominance relations and to condition the generative model toward high-performing regions of the search space. We also show that this non-dominance CPE implicitly estimates the probability of hypervolume improvement (PHVI). To incorporate subjective trade-offs, A-GPS introduces preference direction vectors that encode user-specified preferences in objective space. At each iteration, the model is updated using both Pareto membership and alignment with these preference directions, producing an amortized generative model capable of sampling across the Pareto front without retraining. The result is a simple yet powerful approach that achieves high-quality Pareto set approximations, avoids explicit hypervolume computation, and flexibly captures user preferences. Empirical results on synthetic benchmarks and protein design tasks demonstrate strong sample efficiency and effective preference incorporation.