Skip to content
AI.info

Research

Tab-PET: Graph-Based Positional Encodings for Tabular Transformers

Overview Research area: Machine learning for tabular data — specifically transformer architectures, positional encodings (PEs), and graph-based structure learning. Technical level: Advanced. The paper

Tab-PET: Graph-Based Positional Encodings for Tabular Transformers
arXiv
2511.13338
Published
2025-11-17
Authors
Yunze Leng, Rohan Ghosh, Mehul Motani

AI summary

Overview

  • Research area: Machine learning for tabular data — specifically transformer architectures, positional encodings (PEs), and graph-based structure learning.
  • Technical level: Advanced. The paper includes formal definitions (effective rank), two theorems with proofs deferred to an appendix, graph Laplacian spectral methods, and a large multi-dataset empirical study.
  • Scope (1 sentence): The paper proposes Tab-PET, a framework that estimates a feature-wise graph from tabular data and derives graph Laplacian eigenvector positional encodings for tabular transformers, showing theoretically that PEs reduce the effective rank of learned representations and empirically that they improve performance across 50 classification and regression datasets.

What This Paper Is About

Tabular datasets are small, high-dimensional, and heterogeneous (categorical plus continuous features), and unlike images or text they have no inherent spatial or sequential structure — so transformers, which rely on positional information to guide self-attention, operate on what is essentially an unordered set of features. The prevailing view in the literature (the authors cite Somepalli et al. 2021) is that positional encodings cannot help tabular transformers precisely because tabular data has no structure. This paper challenges that view by showing that a structure can be estimated from the data itself: the authors build a graph over features using association or causality measures, extract Laplacian eigenvectors as positional encodings, and demonstrate both theoretically and empirically that these encodings improve generalization for tabular transformers.

Key Contributions

  1. Tab-PET method. A principled pipeline for constructing positional encodings in tabular domains: (1) learn a feature graph capturing inter-feature associations, using either a causality-based or an association-based approach, and (2) generate feature-wise encodings from the Laplacian eigenvectors of that graph and concatenate them with the input embeddings of the transformer.
  2. Theoretical motivation via rank analysis. A proof that positional encodings give the model the ability to reduce the effective rank (a form of intrinsic dimensionality) of embeddings inside transformer architectures — and that this reduction is stronger when the encodings are aligned with the actual structure of the data. Empirical tests confirm the finding.
  3. Empirical evaluation across benchmarks. Application of Tab-PET to leading tabular transformers — TabTransformer, SAINT, and FT-Transformer (referred to collectively as "3T") — with consistent improvements reported across 50 classification and regression datasets.
  4. Ablations and performance analysis. Ablation studies, statistical significance testing, and comparisons against learnable positional encodings are used to justify the design choices.

Main Findings

  • PEs reduce effective rank. The paper defines effective rank as the exponential of the Shannon entropy of the normalized singular values of the CLS embedding matrix, and shows in Theorem 1 (i.i.d. inputs) that under a single-layer, single-head FT-Transformer, the effective rank is bounded by a quantity that simplifies to approximately r_eff ≈ 1 + d / C_α, where C_α = exp((ατ − 2 c_K c_Q c_q) / √d_T). Larger scaling α of the positional encodings reduces effective rank, but only when τ > 0; with no PEs, τ = 0 and the effective rank can be substantially larger.
  • Alignment with data structure matters. Theorem 2 handles structured inputs (d even, with x_i = θ for i ≤ d/2 and x_i = θ' for i > d/2). With random positional encodings the bound simplifies to approximately r_eff ≈ 1 + d/(2C_α) when C_α ≫ d; when the same PE is shared within each group of similar dimensions, the bound simplifies to r_eff ≈ 1 + 1/C_α for large C_α. Sharing encodings across similar dimensions therefore reduces effective rank much more aggressively.
  • A stated limitation. The authors state that tasks which intrinsically require larger effective rank to be addressed appropriately may not benefit from including PEs.
  • Synthetic experiments confirm the structure hypothesis. On synthetic regression data with input dimensionality fixed at d = 30 and Spearman-based graph PEs, datasets with stronger internal associations benefited more as the positional signal was amplified by higher α. The three regimes were high structure (k ≤ 8), moderate structure (10 ≤ k ≤ 22), and low structure (k > 22). Even in highly unstructured settings, small but consistent gains were observed, attributed to the effective-rank reduction of Theorem 1. Increasing α to 10 degraded performance.
  • Association-based graphs beat causality-based graphs. Averaged over 50 datasets with FT-Transformer as the backbone, Spearman correlation gave the highest average improvement (1.72% classification, 4.34% regression), followed closely by Pearson (1.61%, 4.16%). LiNGAM gave 1.41% / 3.97%, NOTEARS 1.36% / 3.64%, and Chow-Liu 1.17% / 4.29%. Spearman was also the most consistent, rarely degrading performance; Chow-Liu did not perform well on classification.
  • Computational overhead is small for association methods. Average extra time per dataset for graph estimation and PE creation: NOTEARS 76.83 minutes, LiNGAM 10.96 minutes, Pearson 0.78 minutes, Spearman 0.79 minutes, Chow-Liu 0.38 minutes.
  • Graph entropy explains the gap. Causal discovery methods (NOTEARS, LiNGAM) produced graphs concentrated in the low-entropy region — sparse, highly constrained graphs — while Spearman and Pearson produced high-entropy, denser graphs. The denser structures aligned with the strongest performance gains, suggesting PEs carry more useful information when derived from dense feature dependencies than from sparse causal structures.
  • Tab-PET improves all three transformer backbones. On classification, mean rank improved from 4.44 to 2.44 for FT-Transformer, 4.52 to 3.28 for SAINT, and 7.33 to 5.33 for TabTransformer; the baselines were XGBoost at 3.40 and CatBoost at 3.76. On regression, mean rank improved from 4.08 to 2.88 for FT-Transformer, 3.64 to 2.84 for SAINT, and 7.14 to 5.71 for TabTransformer; XGBoost ranked 5.20 and CatBoost 2.96. Lower rank is better.
  • Fixed PEs outperform learnable PEs. The authors adopt a fixed strategy (first and last k Laplacian eigenvectors) and report that fixed PEs empirically outperform learnable PEs built from scratch on tabular datasets. The fixed choice also avoids increasing parametric complexity and keeps comparisons fair across architectures.

Methodology in Plain English

The pipeline has four steps, illustrated in Figure 1 of the paper.

1. Data preparation. Every categorical variable is one-hot encoded, which removes any implicit ordering bias in the raw representation but increases feature dimensionality and therefore the size of the graph built later. Continuous variables are normalized to zero mean and unit variance.

2. Graph estimation. Each processed feature dimension becomes a node in a graph; edges represent statistical or causal dependence. Two families are explored. Causality-based methods assume a linear structural causal model x = Ax + ε and learn a directed acyclic graph using LiNGAM (which relies on non-Gaussianity assumptions) or NOTEARS (which turns structure learning into continuous optimization with smooth acyclicity penalties). Association-based methods set the edge weight w_ij = ρ(x_i, x_j), where ρ is Pearson correlation, Spearman rank correlation, or mutual information; for MI the Chow-Liu algorithm is used so the result stays a DAG.

3. Positional encoding creation. The adjacency matrix is symmetrized as A_sym = ½(A + Aᵀ), and the graph Laplacian L = D − A_sym is computed from the degree matrix D. The top-k and bottom-k eigenvectors are selected, excluding the first (constant) eigenvector, normalized to zero mean and unit variance across nodes, and concatenated into a matrix P = [e₂, …, e_{k+1}, e_{d−k+1}, …, e_d]. These are scaled by a hyperparameter α so P' = α · P. For categorical features spread across several one-hot nodes, the individual encodings are averaged into one consolidated vector per feature. The number k is chosen automatically by a spectral gap analysis that thresholds normalized eigenvalues from both sides.

4. Integration. Each feature's embedding z_i is concatenated with its scaled encoding to give z'_i = [z_i; p'_i] ∈ R^{n+2k}, and these augmented embeddings are fed into the self-attention layers.

Synthetic validation. A synthetic data generator (Algorithm 1) partitions d features into k disjoint groups. Each group has a latent variable θ_g ~ U(−2, 2), each feature is x_f = θ_g · w_f + ε_f with w_f ~ U(−1, 1) and ε_f ~ N(0, 0.01), and the target is a linear function of one fixed group's latent variable: y = w_t · θ_{g*} + b with w_t, b ~ U(−1, 1). Lower k means more structure; higher k means most features sit in their own group and are effectively independent.

Real-data setup. Experiments use 50 tabular datasets from OpenML (25 classification, 25 regression) spanning varied sample sizes, dimensionalities, and categorical ratios. Classification uses stratified sampling to preserve exact class distributions, with a 60:20:20 train-validation-test split; Table 2 notes results are based on five-fold cross-validation. Metrics are balanced accuracy for classification and RMSE for regression, each averaged over five random seeds with early stopping. The α hyperparameter is selected from a set of 9 values between 0.05 and 10 via a greedy search on the validation set. Baselines are XGBoost and CatBoost (with Optuna hyperparameter optimization, 100 trials per dataset) plus TabTransformer, SAINT, and FT-Transformer. TabTransformer is evaluated only on datasets with multiple categorical variables, because its embedding layer — and hence its PEs — applies only to categorical features. Experiments ran on 3 NVIDIA A100 SXM4 80GB GPUs and 3 NVIDIA RTX 6000 Ada Generation GPUs.

Why This Matters

Impact on research. The paper directly contradicts the standing assumption that positional encodings are useless for tabular transformers because tabular data is structure-less. It reframes PEs not as a way to encode known order, but as a mechanism for controlling the dimensionality of the learning problem — a functional role rather than a representational one. The effective-rank framing connects tabular representation learning to a broader literature linking low intrinsic dimensionality with better generalization, and it offers a concrete recipe (spectral graph encodings) that other tabular architectures could adopt.

Real-world applications (tabular data domains the paper identifies as spanning healthcare, finance, and recommender systems):

  • Healthcare — clinical prediction and risk scoring from patient records, which typically contain a few hundred to a few thousand samples with mixed categorical and continuous variables.
  • Finance — credit scoring, fraud detection, and risk modeling, where datasets are similarly small and heterogeneous.
  • Recommender systems — user-item feature tables where feature-order structure is arbitrary.
  • Spreadsheet-style business analytics — the paper contrasts its setting (only feature names and raw values, no external linguistic context) with language-guided models such as TAPAS that exploit column headers and metadata, positioning Tab-PET for the case where such semantic cues are unavailable.

Industry relevance. On classification, Tab-PET variants occupy the top two mean ranks (FT-Transformer +PET at 2.44 and SAINT +PET at 3.28), ahead of XGBoost (3.40) and CatBoost (3.76). On regression, CatBoost still ranks first (2.96) with SAINT +PET (2.84) and FT-Transformer +PET (2.88) close behind, so the method narrows the gap between transformers and gradient-boosted trees that dominate industrial tabular work. The cost of the best-performing association approach is modest — Spearman adds only 0.79 minutes of graph estimation and PE creation on average per dataset. The method requires no external metadata and no increase in parametric complexity, making it a relatively low-friction addition to existing tabular transformer pipelines.

Future Directions

  • When PEs do not help. The authors explicitly note that tasks intrinsically requiring a larger effective rank may not benefit from PEs. Characterizing which datasets fall into that category — and detecting them before training — remains open.
  • Improving learnable PEs. Fixed spectral encodings outperformed learnable PEs built from scratch in this study. Whether better-designed learnable schemes (the paper mentions elastic PEs and eigenvector weighting from prior work) could close or reverse this gap on tabular data is unanswered.
  • Scaling graph estimation. NOTEARS and LiNGAM introduce tunable hyperparameters and, in the case of NOTEARS, substantial cost (76.83 minutes on average versus 0.38–0.79 for association methods). Extending the approach to larger feature spaces without prohibitive overhead is a practical open problem.
  • Deeper structural analysis. The graph-entropy result explains why denser association graphs outperform sparse causal ones, but the paper's own analysis section (7.4) is truncated in the provided content and points to additional comparison work in the technical appendix. The relationship between graph density, eigenvector choice (k selection), and downstream gains is a natural avenue for further study.

Target Audience

This paper is best suited to machine learning researchers and graduate students working on tabular deep learning, tabular foundation models, or transformer architectures who want to understand how inductive biases can be injected into data with no native structure. It will also interest practitioners applying transformers to small, heterogeneous datasets — particularly those comparing deep models against gradient-boosted trees — since it reports both theoretical motivation and a broad multi-dataset empirical comparison. Readers will need comfort with linear algebra (graph Laplacians, singular value decompositions, spectral methods) and transformer internals; the theoretical sections are not introductory. Those without that background can still follow the methodology, synthetic experiments, and results tables, but should note that the appendices containing detailed dataset properties, preprocessing steps, hyperparameter settings, proofs, and per-dataset results are referenced but not included in the provided paper content, and the main-results discussion in Section 7.4 is truncated mid-sentence.

Authors’ abstract

Supervised learning with tabular data presents unique challenges, including low data sizes, the absence of structural cues, and heterogeneous features spanning both categorical and continuous domains. Unlike vision and language tasks, where models can exploit inductive biases in the data, tabular data lacks inherent positional structure, hindering the effectiveness of self-attention mechanisms. While recent transformer-based models like TabTransformer, SAINT, and FT-Transformer (which we refer to as 3T) have shown promise on tabular data, they typically operate without leveraging structural cues such as positional encodings (PEs), as no prior structural information is usually available. In this work, we find both theoretically and empirically that structural cues, specifically PEs can be a useful tool to improve generalization performance for tabular transformers. We find that PEs impart the ability to reduce the effective rank (a form of intrinsic dimensionality) of the features, effectively simplifying the task by reducing the dimensionality of the problem, yielding improved generalization. To that end, we propose Tab-PET (PEs for Tabular Transformers), a graph-based framework for estimating and inculcating PEs into embeddings. Inspired by approaches that derive PEs from graph topology, we explore two paradigms for graph estimation: association-based and causality-based. We empirically demonstrate that graph-derived PEs significantly improve performance across 50 classification and regression datasets for 3T. Notably, association-based graphs consistently yield more stable and pronounced gains compared to causality-driven ones. Our work highlights an unexpected role of PEs in tabular transformers, revealing how they can be harnessed to improve generalization.

Read the original paper