Skip to content
AI.info

Research

Distributionally Robust Feature Selection

Overview Research area: Machine learning — feature selection, distributionally robust optimization (DRO), and group-robustness across subpopulations. Technical level: Advanced (the problem framing is

Distributionally Robust Feature Selection
arXiv
2510.21113
Published
2025-10-24
Authors
Maitreyi Swaroop, Tamar Krishnamurti, Bryan Wilder

AI summary

Overview

Research area: Machine learning — feature selection, distributionally robust optimization (DRO), and group-robustness across subpopulations. Technical level: Advanced (the problem framing is accessible; the method relies on Bayesian conditional expectations, Gaussian kernel smoothing, and gradient estimation). Scope: One-sentence scope: the paper proposes a model-agnostic, noise-based continuous relaxation for selecting a fixed subset of k features that supports high-performing downstream models simultaneously across multiple populations.

What This Paper Is About

Collecting features is often expensive — it may require adding survey items, sensors, or clinical questions — so practitioners use pilot data on many features to decide which few to keep. The paper's core problem is choosing a limited set of features that work well not just on average, but for every subpopulation (e.g., clinics, states, or demographic groups), even when those populations have conflicting or shifting data distributions. The goal is a single global feature subset that any downstream model can later use, without committing to a particular model architecture during selection.

Key Contributions

  1. A model-agnostic distributionally robust feature selection method that selects a subset of features minimizing the maximum expected loss across a set of candidate distributions, without requiring assumptions about the architecture or differentiability of the downstream predictive model.
  2. A continuous relaxation of the combinatorial selection problem using parameters that govern a noise-injection process over each covariate, creating a differentiable measure of each feature's utility and enabling tractable gradient-based optimization.
  3. A closed-form, kernel-based objective: the Bayes-optimal formulation is reduced to a Gaussian kernel smoothing expression, letting the method be solved with standard stochastic gradient descent while fitting each population's model only once.
  4. Empirical validation on synthetic and real-world data, comparing against naive selection strategies and DRO baselines adapted for feature selection.

Main Findings

  • Synthetic linear dataset (populations A, B, C; 15 features): With a budget of 5 features, the method achieves consistently low error across all populations, performing on par with DRO-XGBoost and DRO-Lasso and outperforming both in population A. At budget 10, overall performance improves and the method maintains lower MSE than other methods, including DRO-Lasso.
  • Conflicting-coefficient failure of naive methods: Because features X0–X4 have strong but sign-reversed effects in populations A and B, vanilla LASSO performs best on population C (whose relevant features do not conflict) and poorly on A and B despite the linear setting; vanilla XGBoost shows the same trend.
  • Synthetic nonlinear dataset (populations A, B, C, D; 50 features; budget 8): The nonlinear setting disadvantages vanilla Lasso and, especially, DRO-Lasso (worst performer). The Embedded MLP baseline is imbalanced against population C. The proposed method is comparable to the best-performing baselines (DRO and vanilla XGBoost) at both budgets.
  • Low variance: The method shows consistently low variance across both synthetic experiments, and the relative ranking of feature selection methods is consistent across the choice of downstream prediction model.
  • ACS income regression (CA, FL, NY): The method achieves an order-of-magnitude lower MSE and substantially higher R² than all baselines across all populations, with consistently low run-to-run variance. DRO-XGBoost shows significantly higher error and variance, and Lasso variants outperform XGBoost on this dataset.
  • UCI Adult Income classification (Female and Male populations): The method achieves the highest accuracy in both groups while maintaining comparable log loss and generally lower run-to-run variance. XGBoost outperforms it on log loss for the Female population, but the proposed method is more balanced across the two populations.
  • Exact numeric metric values are not reported in the main text — they are deferred to Tables 1–4 in Appendix D, which are not included in the provided content.

Methodology in Plain English

The problem is formalized as choosing a binary mask over features of size exactly k, minimizing the worst-case loss of the best model trained on the masked covariates, given samples from each population.

Relaxation via noise. Because the binary mask makes the problem combinatorial (NP-hard even for linear models and one population), the mask is relaxed to continuous values. Naively scaling each input by its weight fails, because a flexible model can simply undo the scaling. Instead, each parameter controls the variance of Gaussian noise added to that covariate: the observed variable is S(α) = X + ε(α), where ε(α) ~ N(0, diag(α)). A parameter of 0 means the feature is clean; a parameter approaching infinity means the observation carries no information about that covariate. A regularization term (for example, 1/‖α‖₁) controls sparsity, with λ tuned to reach the desired cardinality.

Targeting the Bayes-optimal predictor. Rather than differentiating through the training of a specific model (which would be expensive, unstable, and tied to one model class such as neural networks), the objective targets the loss of the Bayes-optimal predictor E[Y | S(α)]. Under mean squared error, this loss equals the conditional variance of Y given S(α), and applying the law of total variance yields a population-level objective depending only on the conditional variance of μᵢ(X) = E[Y | X] given S(α) (Theorem 1, proof in Section A.1).

Closed-form kernel objective. Each population's conditional mean function μᵢ is estimated just once by fitting any suitable model on that population's samples. Bayes' theorem under the Gaussian noise model then converts the empirical objective into a Gaussian kernel-weighted average of the fitted μᵢ values at the observed samples, where the noise parameters α set the kernel bandwidth (Theorem 2, proof in Section A.2).

Optimization. Gradient descent is applied to the negative of this objective, with the outer expectation approximated by b Monte Carlo samples of S drawn using the reparameterization trick S = X + √α ⊙ ε, so gradients flow through both the kernel weights and the sampled values. The inner sum is restricted to the k-nearest neighbors of each sample for efficiency. The worst-case over populations is handled either by taking the maximum loss or a temperature-controlled softmax over population losses. Finally, the k features with the smallest optimized noise parameters are selected. Algorithm details and computational complexity are given in Appendix B.

Experimental setup. Baselines include Lasso (linear/logistic regression) with features chosen by coefficient magnitude along the regularization path, XGBoost with internal feature importance, DRO variants of both, and an Embedded MLP with a learnable mask trained via DRO — all trained on a pooled dataset of all populations. Each dataset is split 60:40 per population into a feature-selection dataset and downstream datasets, with the latter split 80:20 for training and evaluation. Downstream models (random forest and MLP) are retrained independently for each selection method and evaluated with MSE (regression) and Log Loss (classification), plus R² and accuracy on real data. The method is implemented in PyTorch and downstream models in scikit-learn.

Why This Matters

Impact on research: The paper identifies and fills a gap at the intersection of two literatures — feature selection (which usually targets a single distribution) and DRO (which usually trains a single robust model rather than selecting robust features). It revives noise-injection-based variable selection, originally proposed by Grandvalet (2000), in a distributionally robust setting, and shows that targeting the Bayes-optimal predictor avoids backpropagating through model training — which matters because common tabular models such as decision trees and random forests are not readily differentiable.

Real-world applications:

  • Healthcare screeners: reducing long instruments such as the 26-item PHQ (Spitzer et al., 1999) to short forms — the paper notes the PHQ is typically cut to standard 9- or 2-question versions (Kroenke et al., 2001; Kroenke et al., 2003) — while keeping performance stable across patient groups.
  • Health system deployment: choosing survey items from a pilot study that will be administered across diverse clinics and populations.
  • Public health and census-style prediction: the ACS experiments target household income prediction across California, Florida, and New York; the UCI Adult experiments target income prediction across sex groups.
  • Sensor or data-collection budgeting: applications where each physical sensor or data field carries an operational cost, requiring a fixed feature set for all users.

Industry relevance: Any organization that must fix a small, standardized feature set for all users — insurers,医院的 clinical risk prediction teams, survey platforms, and teams building tabular models where random forests or gradient boosting are preferred over neural networks — benefits from a selection procedure that is model-agnostic, does not require differentiable models, and explicitly guards against neglecting a subpopulation.

Future Directions

  • Scaling and complexity: the paper defers computational complexity to Appendix B; extending the approach to very high-dimensional covariate pools remains a natural step, especially given the k-nearest-neighbor approximation used for the inner sum.
  • Beyond known, discrete groups: the formulation assumes a specified set of populations 𝒫; handling unknown or continuously parameterized shift families, or populations not observed during selection, is open.
  • Theoretical guarantees: the paper provides equivalence theorems (Theorems 1 and 2) but does not report finite-sample bounds connecting the relaxed objective to the original ℓ₀-constrained problem with a zero-one mask.
  • Cardinality control and downstream validation: the regularization parameter λ is varied to obtain a solution with the desired cardinality; whether the selected set matches the true informative features, and how Bayes-optimal proxies translate to finite-data downstream performance, are questions raised by the approach.

Target Audience

Researchers and practitioners working on robust machine learning, feature selection, and subgroup fairness, as well as applied teams in healthcare, public policy, and survey design who must commit to a small fixed set of features before knowing which models will be deployed. Readers need comfort with probability, optimization, and DRO concepts; the core idea is understandable without the proofs.

Authors’ abstract

We study the problem of selecting limited features to observe such that models trained on them can perform well simultaneously across multiple subpopulations. This problem has applications in settings where collecting each feature is costly, e.g. requiring adding survey questions or physical sensors, and we must be able to use the selected features to create high-quality downstream models for different populations. Our method frames the problem as a continuous relaxation of traditional variable selection using a noising mechanism, without requiring backpropagation through model training processes. By optimizing over the variance of a Bayes-optimal predictor, we develop a model-agnostic framework that balances overall performance of downstream prediction across populations. We validate our approach through experiments on both synthetic datasets and real-world data.

Read the original paper