Skip to content
AI.info

Research

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data

Overview Research area: Statistical graphics and visual analytics for multivariate categorical data, spanning visualization methodology (stat.ML), multivariate categorical analysis, and exploratory da

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data
arXiv
2607.15018
Published
2026-07-16
Authors
Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, ShengLi Tzeng, Han-Ming Wu

AI summary

Overview

  • Research area: Statistical graphics and visual analytics for multivariate categorical data, spanning visualization methodology (stat.ML), multivariate categorical analysis, and exploratory data analysis.
  • Technical level: Intermediate. The paper is written for statistically trained readers: it uses matrix algebra, homogeneity analysis (HOMALS), alternating least squares, and metric-space arguments, but the objectives and outputs are visual rather than predictive.
  • Scope: The paper proposes cGAP, a full-matrix visualization framework that uses a three-dimensional HOMALS embedding to color-code a raw categorical data heatmap and to build subject-level and variable-level proximity matrices, with seriation to expose cluster and outlier structure.

What This Paper Is About

Visualization tools for high-dimensional categorical data lag behind those for continuous data: existing categorical methods either scale poorly, rely on low-dimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. cGAP addresses this by keeping the original observation table as the central visual object while augmenting it with geometric structure derived from HOMALS. Subjects and category levels are embedded in a three-dimensional Euclidean space, the embedding is mapped to red-green-blue coordinates so that similar response patterns receive similar colors, and seriation reorders rows and columns to reveal coherent clusters, outliers, and local-to-global structure.

Key Contributions

  1. A full-matrix visualization framework for high-dimensional categorical data. cGAP preserves the original observation table as a central analytic view rather than reducing the analysis to a low-dimensional scatterplot. It handles nominal, ordinal, and binary data.
  2. An interpretable, HOMALS-based color construction. Rather than assigning colors from variable-specific category labels (which share no global ordering), cGAP derives colors from the geometry of the embedding, so that similar categorical patterns are represented by similar colors. The three embedding axes are ordered by decreasing retained variation (H1, H2, H3) and mapped to the R, G, and B channels respectively.
  3. An integrated three-view display with seriation. A HOMALS-guided heatmap of the raw data matrix is coordinated with a subject proximity matrix and a variable proximity matrix, then reordered using Rank-Two Ellipse (R2E), Rank-One Tree (R1T), and the hybrid HCT-R2E seriation procedures (with average linkage for agglomerative hierarchical clustering tree seriation).
  4. Formal properties linking embedding geometry to the display. The paper derives barycentric traceability (Proposition 1), affine color traceability (Corollary 1), a projection-distortion decomposition (Proposition 2), and a ray-preserving contrast transform (Proposition 3), plus a retained-variation criterion for assessing a three-dimensional embedding.

Main Findings

  • Retained-variation check. The paper defines the proportion of attainable variation retained by a p-dimensional embedding as the ratio of the sum of the first p average discrimination powers to the sum over the full embedding rank. In cGAP this serves as a practical check on whether three dimensions are adequate for color encoding and structural visualization. For the mammalian dentition data, HOMALS fitted with order constraints retained approximately 67% of the attainable variation.
  • Barycentric traceability. Each subject's HOMALS coordinate is exactly the barycenter (average) of the category coordinates the subject selected. Because the RGB relocation map is affine, the same identity holds for displayed colors: the displayed color of a subject is the affine average of the colors of its selected categories. This is what justifies treating colors as "profile colors."
  • Projection distortion is one-sided and exact. Restricting the display to p dimensions can only underestimate full-space pairwise distances between subjects, and the discrepancy is exactly the squared distance contributed by the discarded coordinates. The same decomposition applies to category coordinates.
  • Contrast enhancement preserves direction. The power transformation with contrast parameter q ≥ 1 moves points outward from the cube center. When q = 1 the display is unchanged; as q approaches infinity, points are pushed progressively toward the cube surfaces. Because the multiplicative factor is positive, each point stays on the same ray from the RGB cube center and radial order along a shared ray is preserved, so the transform changes saturation and contrast without altering directional structure inherited from HOMALS.
  • Embedding axes are only identified up to sign and permutation. This affects which absolute hues appear in a figure but leaves pairwise distances, cluster relationships, and block patterns in the sorted matrices unchanged. The paper therefore treats hue relationally, not semantically.
  • Animal-grouping example. Applied to the animal-grouping data of Nishisato (2007) — fifteen students grouping thirty-five animals, producing ninety-five student-group categories — the HCT-R2E-sorted display identifies four main perceptual groups: non-primate mammals (sea green), birds (pink), reptiles and frogs (purple), and primates (green), plus a stand-alone cluster of alligators. The alligator elicited diverse classifications and is rendered as a blend of purple, sea green, and green. The ostrich acts as a bridging animal between the mammal and bird clusters, reflecting its categorization with mammals by five students (S1, S7, S11, S14, S15). The monkey and chimpanzee share nearly identical color profiles and high proximity. Students S10 and S11 show response patterns that differ from most other students, and additional local anomalies include the treatment of pigeon by Student S12 and of sparrow and cow by Student S10. Category colors obey the relativity principle: Student S5's sixth category and Student S8's fifth category both contain birds, yield identical HOMALS solutions, and are displayed in the same pink hue.
  • Mammalian dentition example. For 66 mammals described by eight ordinal dentition variables, the variable proximity view shows that top and bottom versions of the same tooth types lie close together, especially for molars and canines. Category colors reveal an opposition: mammals with fewer molars tend to retain canines, while mammals with more molars often do not. Carnivora form coherent groups; Rodentia- and Lagomorpha-related mammals occupy a separate region with more molars and fewer canines; Artiodactyla are grouped by the joint absence of incisors and canines. The armadillo appears as an extreme outlying configuration at one end of the dentition space, and the Common Mole appears closer to a carnivore-like profile than its taxonomic label alone would suggest. Overall, the sorted heatmap suggests a gradient from carnivore-like dentition (retained canines, reduced molar counts) toward herbivore- and rodent-like dentition.
  • Mushroom and COG examples. The mushroom dataset from the UCI Machine Learning Repository comprises 8,124 samples from the Agaricus and Lepiota families described by 22 nominal physical-attribute variables plus edibility information; the paper states it is used to test whether cGAP can expose interpretable nominal structure in a large benchmark. The COG profiles example is a 2,296 × 5,061 binary dataset. Table 2 characterizes the roles and expected outputs of these examples (locally decisive variables, edibility regimes, and informative missingness for mushrooms; phylogenetic blocks, complementary present/absent modules, and multiscale genomic structure for COGs). The provided text of Section 4.2 is truncated mid-sentence, so the specific mushroom and COG results are not reported in the available content.

Methodology in Plain English

The starting point is a table of N subjects measured on J categorical variables, where variable j has c_j categories. Each variable is turned into an indicator matrix that marks which category each subject selected. HOMALS (homogeneity analysis) then searches simultaneously for coordinates for subjects and for category levels so that subjects sit near the categories they chose and categories sit near the subjects who chose them. This is done by minimizing an average reconstruction loss, solved by alternating least squares: given subject coordinates, category coordinates are updated by least squares; given category coordinates, subject coordinates are updated as the average of the quantified categories each subject selected; the process repeats until convergence. Centering and normalization constraints prevent the trivial all-zero solution.

Because the resulting subject and category points live in a common Euclidean space, cGAP can turn them into colors. Coordinates are linearly relocated into the unit cube and read off as red, green, and blue intensities. Points near the center become grayish; extreme configurations move toward the cube faces and corners. Edges connecting subjects to categories inherit the category color, linking the embedding back to the original matrix entries. If a few outlying configurations pull everything else toward the gray center and dull the contrast, a power transformation with parameter q pushes points outward along the same rays from the cube center to restore visual separation.

Proximity is handled in two ways. For subjects, the distance is plain Euclidean distance in the three-dimensional embedding. For variables, the paper defines a weighted dissimilarity that sums, over subjects, the embedding distance between the categories the two variables assigned to each subject. The paper verifies this function satisfies non-negativity, identity of indiscernibles, symmetry, and the triangle inequality, while cautioning that a zero value means the two variables are indistinguishable under this induced comparison for the observed data, not that they are semantically identical. Finally, seriation algorithms reorder the rows and columns of the heatmap and of both proximity matrices so that similar subjects and variables appear adjacent, making blocks, outliers, and local-to-global structure visible. The paper demonstrates the whole pipeline on small illustrative data (fifteen students, thirty-five animals) and then on larger ordinal, nominal, and binary datasets.

Why This Matters

Impact on research. The paper targets a genuine gap: matrix-based visualization such as data images and generalized association plots works well for continuous data because all variables share one quantitative color scale, but categorical variables have no shared scale. cGAP supplies a principled way to construct one from an embedding, and its formal properties (traceability, distortion, contrast preservation) let analysts reason about how much of the original structure the display preserves. Keeping the raw data matrix visible alongside the derived views means conclusions remain traceable to individual observations rather than to an abstract low-dimensional plot.

Real-world applications:

  • Genomics screening: the paper applies cGAP to 2,296 × 5,061 binary presence/absence profiles from the Clusters of Orthologous Genes database, where the goal is large-scale exploratory screening for phylogenetic blocks and complementary present/absent modules.
  • Biomedical and morphological data: the mammalian dentition example uses 66 mammals and eight ordinal tooth-count variables, revealing anatomical symmetry between upper and lower jaws, functional gradients, and taxonomy-aligned exceptions such as the armadillo.
  • Benchmark datasets with class labels: the mushroom example uses 8,124 samples and 22 nominal attributes plus edibility, testing whether nominal structure and class-related patterns can be read visually.
  • Social and behavioral surveys: the animal-grouping study, in which fifteen students classified thirty-five animals, illustrates how cGAP renders respondent-level differences and ambiguous items (the alligator, the ostrich) as blended colors and bridging positions.

Industry relevance. Any organization holding wide categorical tables — survey responses, clinical coding, product attributes, taxonomies, sensor states — needs tools that scale to many variables and many categories without collapsing the data into an unreadable projection. cGAP's emphasis on keeping the original matrix, on color as relative similarity rather than fixed semantics, and on coordinated sorted views, gives analysts a screening environment that complements predictive modeling. Its user-specified contrast parameter provides a simple knob for making minority configurations visible in large tables.

Future Directions

  • Perceptually optimized and accessible palettes. The paper states explicitly that a more comprehensive study of perceptually optimized palettes and color-vision-deficiency-safe encodings remains an important direction for future work, since RGB is adopted as a display-native coordinate system rather than a claim of perceptual optimality.
  • Choosing the number of dimensions and the contrast parameter. The retained-variation criterion is offered as a practical check on whether three dimensions suffice, but the paper does not resolve how many dimensions to use in general, nor how to select the contrast parameter q in a principled way; the transformation matrix, the object central to these extensions, is introduced but the truncated text does not develop them.
  • Handling axis sign and permutation ambiguity. Because HOMALS axes are unique only up to sign changes and axis permutation, specific hues are not stable across runs. The paper's convention of ordering axes by retained variation mitigates this for hue ordering, but a systematic convention or reporting standard for absolute color output is left open.
  • Scaling to and interpreting very large categorical tables. The COG example (2,296 × 5,061) is framed as large-scale exploratory screening. Whether the coordinated three-view display and seriation remain legible as the number of variables and categories grows further, and what computational or interface strategies would be needed, are questions the paper's concluding section takes up but whose content is not available in the provided text.

Target Audience

This paper is most useful to statisticians and data scientists working with multivariate categorical data — nominal, ordinal, or binary — who need exploratory visualization that does not discard the original observation table. It will also interest researchers in statistical graphics and visual analytics who follow matrix visualization, seriation, and correspondence-analysis-style methods, as well as applied analysts in genomics, biomedicine, and the social sciences dealing with wide categorical tables. Readers should be comfortable with basic matrix notation, Euclidean distance, and clustering terminology; the paper supplies its own background on HOMALS, so prior expertise in homogeneity analysis is not strictly required. Those looking for predictive benchmarks or model accuracy comparisons will not find them here, since the paper is explicitly framed around interpretability and exploration rather than prediction.

Authors’ abstract

High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on low-dimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. To address this gap, we introduce categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that preserves the original data matrix while augmenting it with interpretable geometric structure. cGAP uses Homogeneity Analysis (HOMALS) to embed subjects and category levels in a three-dimensional Euclidean space and maps the embedding to red-green-blue coordinates so that similar patterns receive similar colors. The framework integrates three coordinated views: a HOMALS-guided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. Seriation algorithms are then used to reorder rows and columns to reveal coherent clusters, outliers, and local-to-global structure. We also derive barycentric traceability, projection-distortion, and contrast-preservation properties that clarify how embedding geometry is transferred to the display. We demonstrate the versatility of cGAP through applications to student-animal classification data, mammalian dentition profiles, mushroom records from the UCI Machine Learning Repository, and the Clusters of Orthologous Genes database. These examples show that cGAP supports transparent exploratory analysis by maintaining traceability between derived visual structure and the original categorical observations. cGAP provides a full-matrix, heatmap-based visualization environment for investigating complex categorical datasets across scientific domains.

Read the original paper