Skip to content
AI.info

Research

From Black-Box Tuning to Guided Optimization via Hyperparameters Interaction Analysis

Overview Research area: Machine learning / AutoML — specifically hyperparameter optimization (HPO), meta-learning, and explainable AI (XAI) via Shapley value analysis. Technical level: Advanced. The p

From Black-Box Tuning to Guided Optimization via Hyperparameters Interaction Analysis
arXiv
2512.19246
Published
2025-12-22
Authors
Moncef Garouani, Ayah Barhrhouj

AI summary

Overview

Research area: Machine learning / AutoML — specifically hyperparameter optimization (HPO), meta-learning, and explainable AI (XAI) via Shapley value analysis.

Technical level: Advanced. The paper assumes familiarity with Bayesian optimization, meta-features, surrogate models, and cooperative game theory (Shapley values).

Scope: The paper proposes MetaSHAP, a framework that predicts which hyperparameters matter for a new dataset-algorithm pair — and in which value ranges — by mining a knowledge base of over 09 million previously evaluated machine learning pipelines, then uses those predictions to guide Bayesian optimization.

What This Paper Is About

Automated hyperparameter tuning methods such as Bayesian Optimization, Tree-structured Parzen Estimators, and evolutionary algorithms search the hyperparameter space blindly, with no prior knowledge of which hyperparameters actually matter or how they interact. This wastes computation and gives practitioners no insight into why one configuration beats another. MetaSHAP's goal is to answer, before any expensive tuning begins, three practical questions: which hyperparameters are likely to be important for this dataset, how they interact, and in which value ranges their influence concentrates.

Key Contributions

  1. A formalization of global, dataset-aware informed hyperparameter tunability grounded in game theory. The paper defines a mapping from a new dataset and algorithm to a set of estimated importance scores, framing similarity-based knowledge transfer as a formal problem rather than an ad hoc heuristic.

  2. A scalable architecture combining meta-learning-based retrieval, surrogate modeling, and SHAP-based analysis over a benchmark of more than 09 million evaluated ML pipelines, spanning 164 classification datasets and 14 classifiers.

  3. A method for extracting actionable tuning insights — not just importance rankings, but recommended value ranges for the most influential hyperparameters, derived from sorted and smoothed SHAP value distributions (e.g., narrowing XGBoost's learning_rate to [0.02, 0.5]).

  4. An empirical validation showing MetaSHAP-guided Bayesian optimization is competitive with, and typically faster than, standard unguided BO across 06 datasets, while reducing search dimensionality by over 50–70%.

Main Findings

  • Faster convergence across nearly all tested datasets: MetaSHAP-guided BO reaches high-performing configurations significantly faster than unguided BO. On the ring dataset, guided BO achieves near-optimal accuracy (above 0.97) from the first iteration, whereas standard BO requires around 12 iterations to reach a similar level.
  • Dramatic gains on titanic: the guided version stabilizes around its optimal accuracy by iteration 3, while standard BO takes approximately 9 iterations — more than three times as long — to catch up.
  • A notable exception on cars: MetaSHAP-guided BO finds a better configuration as early as iteration 2, but vanilla BO eventually surpasses it around iteration 24. The authors attribute this to insufficiently similar analog datasets in the knowledge base.
  • Robustness when guidance is poor: even in the cars case, MetaSHAP does not severely degrade performance, suggesting degradation is graceful rather than catastrophic.
  • Dimensionality reduction: MetaSHAP typically identifies 3 to 5 hyperparameters as dominant per model-dataset pair, cutting tuning complexity by over 50–70% relative to the full hyperparameter set.
  • Interpretable, human-readable recommendations: for XGBoost's learning_rate, the method consistently suggested narrowing the range to [0.02, 0.5], described as aligning with known best practices.
  • Interaction patterns recovered for SVM: on the "ring" dataset, C, coef0, and gamma tend to show synergistic effects, while degree often contributes negligibly. High values of C consistently lead to strong SVM performance, whereas coef0 shows no clear trend across its value range.
  • Guiding strategy details: the guided variant selects the 2–3 most influential hyperparameters and restricts search to their subspaces, holding other hyperparameters at defaults or marginalizing over a narrow prior range.

Methodology in Plain English

MetaSHAP works in four steps, all performed before running any optimization on a new dataset.

First, it describes the new dataset numerically — size, number of features, class counts, statistical moments, entropy, and the accuracy of simple baseline classifiers ("landmarking"). This produces a compact fingerprint of the dataset.

Second, it finds similar datasets in a large historical knowledge base by computing Euclidean distance between dataset fingerprints and retrieving the closest neighbors. The assumption is that datasets with similar characteristics have similar performance landscapes over hyperparameters.

Third, it trains a surrogate regression model on the historical hyperparameter-performance pairs from those neighboring datasets. This surrogate acts as a cheap approximation of how the algorithm would perform for any given hyperparameter configuration.

Fourth, it applies SHAP (Shapley additive explanations) to the surrogate, treating each hyperparameter as a "player" in a cooperative game and the model's performance as the reward to be fairly split. This yields a marginal contribution score for each hyperparameter, plus pairwise interaction indices capturing synergies and redundancies. To convert scores into actionable advice, the method collects SHAP values alongside the actual hyperparameter values used, sorts them, applies a moving average for smoothing, and extracts the sub-ranges where the absolute smoothed SHAP values are largest — those ranges are flagged as the ones worth tuning.

The knowledge base itself was built by systematically tuning 14 algorithms on 164 supervised classification datasets, evaluating roughly 4000 hyperparameter configurations per algorithm per dataset.

Why This Matters

Impact on research. The paper targets a genuine gap: existing importance estimators such as functional ANOVA attribute variance but do not say in which direction to tune; post-hoc Shapley methods like HyperSHAP explain a completed optimization run but cannot steer one; and meta-learning approaches like Auto-sklearn 2 or ranking-based configuration transfer recommend what to try without explaining why. MetaSHAP combines meta-learning's cross-dataset transferability with Shapley values' theoretical grounding, and adds a concrete procedure for turning attribution scores into search-space restrictions. It also explicitly frames the problem as a global rather than local HPO setting, reusing knowledge across datasets instead of discarding it.

Real-world applications:

  • AutoML platforms could use the method to shrink the search space before launching expensive tuning jobs, saving compute budget and wall-clock time.
  • Data science teams in industry could get human-readable guidance — e.g., "tune these 3 to 5 hyperparameters, and keep learning_rate between 0.02 and 0.5" — instead of a black-box optimizer's output.
  • Resource-constrained settings (small teams, limited GPUs) benefit most from reducing the number of tuned dimensions by over 50–70%.
  • Model debugging and auditing could use the interaction heatmaps to explain why a model behaves as it does, and whether individual hyperparameters have monotone, synergistic, or negligible effects.

Industry relevance. Hyperparameter tuning is one of the most compute-intensive parts of the machine learning lifecycle, and its cost grows with the number of hyperparameters. Any method that reliably prunes the search space without sacrificing final accuracy translates directly into lower cloud costs and faster iteration. The paper's finding that MetaSHAP can act as a complementary layer on top of existing BO implementations "without requiring major algorithmic changes" lowers the barrier to adoption.

Future Directions

  • Handling poor analog retrieval. The cars dataset showed vanilla BO eventually overtaking guided BO, which the authors attribute to the absence of sufficiently similar datasets in the knowledge base. How to detect and gracefully handle such cases — or broaden retrieval — remains open.
  • Generalizing to regression tasks. The authors explicitly name regression as a next target beyond the classification setting studied here.
  • Extending to deep neural networks. The current knowledge base covers 14 classifiers, primarily classical algorithms; transferring the approach to deep networks and their much larger, more structurally different hyperparameter spaces is an open direction.
  • Broader "parameterized learning algorithms." The conclusion frames the general ambition as extending MetaSHAP to other forms of parameterized learning, contributing toward more intelligent and explainable machine learning systems overall.

Target Audience

This paper is most valuable to AutoML and hyperparameter optimization researchers working on meta-learning, transfer learning across tasks, or explainability of optimization processes. It is also relevant to practitioners building or operating AutoML pipelines who need to reduce tuning cost while retaining interpretability, and to applied machine learning engineers who want principled guidance on which hyperparameters to prioritize and in what ranges. Readers without background in Shapley values, meta-features, or Bayesian optimization will find the theoretical sections dense; the results and the tuning-range examples are considerably more accessible.

Authors’ abstract

Hyperparameters tuning is a fundamental, yet computationally expensive, step in optimizing machine learning models. Beyond optimization, understanding the relative importance and interaction of hyperparameters is critical to efficient model development. In this paper, we introduce MetaSHAP, a scalable semi-automated eXplainable AI (XAI) method, that uses meta-learning and Shapley values analysis to provide actionable and dataset-aware tuning insights. MetaSHAP operates over a vast benchmark of over 09 millions evaluated machine learning pipelines, allowing it to produce interpretable importance scores and actionable tuning insights that reveal how much each hyperparameter matters, how it interacts with others and in which value ranges its influence is concentrated. For a given algorithm and dataset, MetaSHAP learns a surrogate performance model from historical configurations, computes hyperparameters interactions using SHAP-based analysis, and derives interpretable tuning ranges from the most influential hyperparameters. This allows practitioners not only to prioritize which hyperparameters to tune, but also to understand their directionality and interactions. We empirically validate MetaSHAP on a diverse benchmark of 164 classification datasets and 14 classifiers, demonstrating that it produces reliable importance rankings and competitive performance when used to guide Bayesian optimization.

Read the original paper