Research
Unsupervised Feature Selection Through Group Discovery
Overview Research area: Unsupervised feature selection and representation learning (machine learning, spectral graph methods, high-dimensional clustering). Technical level: Advanced. The paper assumes

- arXiv
- 2511.09166
- Published
- 2025-11-12
- Authors
- Shira Lifshitz, Ofir Lindenbaum, Gal Mishne, Ron Meir, Hadas Benisty
AI summary
Overview
Research area: Unsupervised feature selection and representation learning (machine learning, spectral graph methods, high-dimensional clustering).
Technical level: Advanced. The paper assumes familiarity with graph Laplacians, spectral methods, Gumbel-Softmax relaxation, and stochastic gates.
Scope: The paper introduces GroupFS, a differentiable end-to-end framework that jointly discovers latent groups of related features and sparsely selects the informative ones, using only unlabeled data, and evaluates it on one synthetic benchmark and nine real-world datasets.
What This Paper Is About
Most unsupervised feature selection methods score each feature on its own, even though genuinely informative signals often come from sets of related features such as neighbouring pixels, coupled brain regions, or correlated financial indicators. Existing group-aware methods require either predefined groups or label supervision, so they cannot be applied when the group structure itself is unknown. The paper's goal is to learn those feature groups and select the useful ones directly from unlabeled data in a single differentiable model.
Key Contributions
- GroupFS framework. The authors introduce GroupFS, described as the first end-to-end, fully differentiable framework for unsupervised feature selection that jointly discovers latent feature groups and selects the informative groups among them.
- No predefined partitions or supervision. GroupFS learns latent feature groups automatically rather than assuming fixed a priori groups or using labels, which the authors argue broadens applicability to unlabeled real-world data.
- Composite smoothness-plus-sparsity objective. The method enforces Laplacian smoothness on both a sample graph and a feature graph, adds an orthogonality penalty on group embeddings, and applies a group-level sparsity regularizer that keeps the representation compact.
- Broad empirical validation. Experiments span synthetic data with known ground-truth groups plus nine real-world image, tabular, and biological datasets, with both quantitative clustering comparisons and a qualitative interpretability study.
Main Findings
- Synthetic group recovery. On a 20-dimensional extension of the two-moons dataset (features 1–5 derived from the first coordinate, 6–10 from the second, 11–20 i.i.d. Gaussian noise), GroupFS isolated the informative groups {1:5} and {6:10} and placed the ten noisy dimensions in unselected clusters across all tested noise levels and correlation strengths.
- Correlation strength helps. Final training loss decreased as the intra-group correlation ρ increased, reaching a minimum at ρ = 1, which the authors attribute to a clearer block structure in the correlation matrices. Results were averaged over 10 random seeds.
- Robustness to noise. The final-loss curve as a function of additive Gaussian noise standard deviation remained essentially flat up to a standard deviation of 0.45, indicating robustness to moderate sample-level noise.
- Group-count sensitivity. The effective number of groups needed to fully separate signal from noise is C = 2 + (d − 10). For C ≤ 2 + (d − 10) the model nearly always achieved RG_sim = 1, TPR = 1 and FDR = 0; C = 2 was an exception because noise features merged with informative ones; for C > 2 + (d − 10), TPR and FDR stayed strong but RG_sim gradually declined as informative clusters were fragmented.
- Fixed-budget clustering (Scenario 1). Using an identical feature budget for all methods, GroupFS ranked first or tied for first on 6 out of 9 datasets, outperforming the next-best method by an average of +3.84%. On two of the remaining three datasets it still ranked in the top three. The authors note Yale and NMNIST were especially challenging, since using all features gave the best result on both.
- Adaptive-budget clustering (Scenario 2). When each method could choose its own feature count from {50, 100, 200, 400} (or {2, 4, 8, 10} for HeartDisease), GroupFS achieved the highest accuracy on 6 out of 9 datasets.
- Interpretable vision groups. On the NMNIST 3–8 subset, GroupFS discovered seven spatially coherent pixel groups. The top-ranked group highlighted regions that differentiate 3s from 8s, such as the upper-left loop of 8s that is typically absent in 3s; lower-ranked groups corresponded mainly to background pixels.
- Interpretable tabular groups. On the UCI Student Performance dataset (395 samples, 30 features), GroupFS found seven interpretable groups; using the top three gave clustering accuracy of 61.3 ± 2.6%. The highest-ranked group covered alcohol consumption, the second covered motivation-related variables (school absences, past academic failures, romantic relationships, intention to pursue higher education), and the third covered parents' education and the mother's job.
Methodology in Plain English
GroupFS builds two similarity graphs: one connecting samples to samples, and one connecting features to features, using a self-tuning kernel that adapts to local density. It then trains a single model with three coordinated objectives.
First, a soft assignment matrix maps each feature to one of C latent groups using the Gumbel-Softmax trick, which lets the model make near-discrete group assignments while staying differentiable. Second, instead of learning a selection gate for every one of the d features, the model learns only C gates, one per group, using the stochastic-gates formulation; each feature's importance is then the weighted average of its group memberships times the group gate values. Multiplying the input by these gates masks out unimportant features. Third, the sample-wise smoothness loss rebuilds the sample graph from the gated data and rewards features that vary smoothly across strongly connected samples; the feature-wise smoothness loss, combined with an orthogonality penalty, encourages features that are close in the feature graph to receive similar group assignments while keeping group embeddings diverse; and the group sparsity loss penalizes both the probability that a group's gate is active and the fraction of features assigned to it, pushing toward fewer and smaller active groups.
The final objective is the sum of these three terms, weighted by λ₁ and λ₂. Group logits are warm-started using spectral clustering on the normalized sample Laplacian with p_main = 0.7, gates are initialized to μ_j = 0.5, and Q is initialized as a random orthonormal matrix; these initial assignments are gradually overwritten during training. Features are selected by ranking groups by mean gate value and keeping the top-ranked ones.
Why This Matters
Impact on research. The paper reframes unsupervised feature selection as a joint grouping-and-selection problem rather than an independent per-feature scoring problem, and shows that group structure can be learned without labels. It provides a differentiable template that combines Gumbel-Softmax grouping, stochastic gating, and dual-graph Laplacian smoothness in one objective, which other researchers can reuse or extend.
Real-world applications.
- Neuroimaging: focusing on task-relevant brain regions to lower acquisition cost and aid domain interpretation.
- Behavioral and survey research: reducing the number of expensive or burdensome questionnaire items while preserving predictive structure (demonstrated on the Student Performance dataset).
- Genomics and clinical profiling: the biomedical benchmarks (ALLAML, Lung500, METABRIC, HeartDisease) involve gene-expression and clinical profiles where selected groups can support biological interpretation.
- Computer vision: the pixel-group results on noisy MNIST variants show spatially coherent, class-relevant regions recovered without labels.
Industry relevance. High-dimensional data with unknown group structure is common in finance, sensor networks, and imaging. Because GroupFS is fully differentiable and selects groups rather than isolated features, it produces compact, interpretable feature sets that can reduce data acquisition, storage, and downstream compute costs while remaining usable as a modular preprocessing step for downstream learning tasks.
Future Directions
- Manifold-aware distances. The authors note that the sample and feature graphs rely on Euclidean distances, which can misrepresent data lying on curved, non-Euclidean manifolds, and propose smooth, differentiable manifold-aware distances as future work.
- Dynamic, condition-adaptive grouping. The current model learns a single global notion of group importance, overlooking condition- or time-dependent relevance; the authors propose a dynamic formulation in which both the groups and their importance can evolve.
- Group-count selection. The paper relies on a heuristic for choosing the number of groups C (described in Appendix D) and shows performance degrades when C is too small or too large, leaving principled automatic selection as an open question.
- Extending to larger-scale settings. MGAGR had to be skipped on NMNIST due to impractical runtime, and scalability of joint grouping under very high feature counts remains a practical question raised by the benchmark design.
Target Audience
Researchers and graduate students in machine learning, signal processing, and computational biology working on dimensionality reduction, unsupervised representation learning, or spectral graph methods. It is also relevant to applied scientists in neuroimaging, genomics, and survey-based behavioural research who need interpretable feature subsets without labels, and to practitioners who want a modular, differentiable feature-selection component for downstream pipelines. Readers without a background in graph Laplacians, Gumbel-Softmax, or stochastic gates will need to consult the cited preliminary work first.
Authors’ abstract
Unsupervised feature selection (FS) is essential for high-dimensional learning tasks where labels are not available. It helps reduce noise, improve generalization, and enhance interpretability. However, most existing unsupervised FS methods evaluate features in isolation, even though informative signals often emerge from groups of related features. For example, adjacent pixels, functionally connected brain regions, or correlated financial indicators tend to act together, making independent evaluation suboptimal. Although some methods attempt to capture group structure, they typically rely on predefined partitions or label supervision, limiting their applicability. We propose GroupFS, an end-to-end, fully differentiable framework that jointly discovers latent feature groups and selects the most informative groups among them, without relying on fixed a priori groups or label supervision. GroupFS enforces Laplacian smoothness on both feature and sample graphs and applies a group sparsity regularizer to learn a compact, structured representation. Across nine benchmarks spanning images, tabular data, and biological datasets, GroupFS consistently outperforms state-of-the-art unsupervised FS in clustering and selects groups of features that align with meaningful patterns.