Skip to content
AI.info

Research

Clustering Approaches for Mixed-Type Data: A Comparative Study

Overview Research area: Unsupervised machine learning / statistical clustering, specifically clustering of datasets that contain both continuous and categorical variables ("mixed-type" data). Technica

Clustering Approaches for Mixed-Type Data: A Comparative Study
arXiv
2511.19755
Published
2025-11-24
Authors
Badih Ghattas, Alvaro Sanchez San-Benito

AI summary

Overview

Research area: Unsupervised machine learning / statistical clustering, specifically clustering of datasets that contain both continuous and categorical variables ("mixed-type" data).

Technical level: Intermediate. The paper assumes familiarity with clustering concepts (distance-based vs. probabilistic models, cluster overlap, evaluation metrics such as the Adjusted Rand Index), but the comparison itself is framed around experimental design factors that are easy to follow.

Scope in one sentence: A comparative simulation study of six existing clustering methods for mixed-type data, examining how their performance changes as the data-generating scenario varies.

What This Paper Is About

Most clustering algorithms are designed for data that is entirely numeric or entirely categorical, and relatively few tools exist for the common case where a dataset mixes both types of variables — for example, a table combining age and income with occupation and region. This paper surveys the state of the art in mixed-type clustering and then pits the leading methods against each other in a controlled simulation study. The goal is to understand which methods hold up under which conditions, rather than to propose a new algorithm.

Key Contributions

  1. A consolidated review of the state of the art in mixed-type clustering, covering both distance-based and probabilistic families of methods.
  2. A systematic simulation-based comparison of six methods — the distance-based approaches k-prototypes, PDQ, and convex k-means, and the probabilistic approaches KAMILA (KAy-means for MIxed LArge data), mixtures of Bayesian networks (MBNs), and latent class models (LCM).
  3. An analysis of how experimental conditions drive performance, varying the number of clusters, degree of cluster overlap, sample size, dimensionality, proportion of continuous variables, and the clusters' distribution.
  4. Identification of a hard regime where no evaluated method succeeds — situations where variables interact strongly and depend explicitly on cluster membership — together with a ranking of the strongest performers (KAMILA, LCM, and k-prototypes) by Adjusted Rand Index, and a note that all methods are available in R.

Main Findings

  • Overlap, continuous-variable share, and sample size matter most: The degree of cluster overlap, the proportion of continuous variables in the dataset, and the sample size all had a significant impact on the performance observed across methods.
  • A difficult regime exists with no good answer: When strong interactions exist between variables and there is an explicit dependence on cluster membership, none of the evaluated methods demonstrated satisfactory performance.
  • Three methods led the pack: KAMILA, LCM, and k-prototypes exhibited the best performance with respect to the Adjusted Rand Index (ARI). The abstract does not specify how these three compare against one another, nor does it report numerical scores.
  • Implementation is accessible: All the compared methods are available in R, so the comparison can be reproduced by practitioners without re-implementing the algorithms.

Methodology in Plain English

The authors did not test the methods on a single benchmark dataset. Instead, they generated synthetic data through a range of simulation models, deliberately varying the conditions under which clustering is attempted: how many clusters are present, how much those clusters overlap, how large the sample is, how many variables there are, how much of the data is continuous versus categorical, and how the clusters are distributed. Each method was then run across this grid of scenarios, and its output was compared against the known cluster labels using the Adjusted Rand Index. This design lets the authors attribute performance differences to specific properties of the data rather than to the quirks of one dataset. The abstract does not describe the number of simulation settings, the specific values chosen for each factor, or the simulation code.

Why This Matters

For researchers, the paper provides a structured map of which method families are worth using under which data conditions — and, importantly, flags a regime where the current generation of methods fails, which is a target for future methodological work. It also makes an implicit point about evaluation practice: a single benchmark comparison can be misleading when performance depends so heavily on overlap and variable composition.

Domains where data is naturally mixed-type and this comparison is relevant include:

  • Healthcare and clinical records, where patient tables mix laboratory measurements with categorical diagnoses, treatment codes, and demographics.
  • Customer and market segmentation, where firms cluster customers using both spending amounts and categorical attributes such as channel, region, or product preference.
  • Social science and survey research, where questionnaires combine Likert-type scales with categorical responses.
  • Operational and industrial telemetry, where sensor readings arrive alongside categorical machine states or fault codes.

For industry, the results are a caution: the choice of clustering algorithm is not a minor implementation detail. A method that performs well on one dataset may degrade sharply as overlap increases or as the mix of continuous and categorical features shifts, so validation on the specific data at hand remains necessary. The availability of all methods in R lowers the barrier to running such comparisons internally.

Future Directions

  • Develop methods for the failure regime: The case of strong variable interactions combined with explicit dependence on cluster membership is left unsolved, which points directly at a gap for new model development.
  • Extend the comparison to real datasets: Simulation gives control over the data-generating process, but the abstract does not indicate whether real-world mixed-type benchmarks were included; validating the simulation conclusions on real data would be a natural follow-up.
  • Broaden the experimental factors: Other dimensions — missing values, noise variables, unequal cluster sizes beyond those simulated, or very high dimensionality — are not mentioned in the abstract and could be explored.
  • Scale and computational cost: The abstract reports performance only in terms of ARI; runtime, memory, and behavior on large datasets are not addressed, which matters especially given that one method is designed for "MIxed LArge" data.

Target Audience

This paper is most useful for applied statisticians and data scientists who regularly cluster datasets with a mix of numeric and categorical variables and need guidance on which method to choose. It also serves machine learning researchers looking for a clear statement of where current mixed-type clustering methods break down, and instructors or students seeking a structured entry point into the literature, since the reviewed methods are all available in R.

Authors’ abstract

Clustering is widely used in unsupervised learning to find homogeneous groups of observations within a dataset. However, clustering mixed-type data remains a challenge, as few existing approaches are suited for this task. This study presents the state-of-the-art of these approaches and compares them using various simulation models. The compared methods include the distance-based approaches k-prototypes, PDQ, and convex k-means, and the probabilistic methods KAy-means for MIxed LArge data (KAMILA), the mixture of Bayesian networks (MBNs), and latent class model (LCM). The aim is to provide insights into the behavior of different methods across a wide range of scenarios by varying some experimental factors such as the number of clusters, cluster overlap, sample size, dimension, proportion of continuous variables in the dataset, and clusters' distribution. The degree of cluster overlap and the proportion of continuous variables in the dataset and the sample size have a significant impact on the observed performances. When strong interactions exist between variables alongside an explicit dependence on cluster membership, none of the evaluated methods demonstrated satisfactory performance. In our experiments KAMILA, LCM, and k-prototypes exhibited the best performance, with respect to the adjusted rand index (ARI). All the methods are available in R.

Read the original paper