Skip to content
AI.info

Research

LATTLE: LLM Attention Transplant for Transfer Learning of Tabular Data Across Disparate Domains

Overview Research area: Transfer learning and representation learning for tabular (spreadsheet/database-style) data, using large language models as a knowledge source. Published on arXiv (arXiv:2511.0

LATTLE: LLM Attention Transplant for Transfer Learning of Tabular Data Across Disparate Domains
arXiv
2511.06161
Published
2025-11-08
Authors
Ibna Kowsar, Kazi F. Akhter, Manar D. Samad

AI summary

Overview

Research area: Transfer learning and representation learning for tabular (spreadsheet/database-style) data, using large language models as a knowledge source. Published on arXiv (arXiv:2511.06161v2 [cs.LG], 23 Jan 2026) by Ibna Kowsar, Kazi F. Akhter, and Manar D. Samad of the Department of Computer Science, Tennessee State University, Nashville, TN, USA.

Technical level: Advanced. The paper assumes familiarity with transformer attention (query/key/value projections), self-attention versus cross-attention, fine-tuning objectives, and standard tabular benchmarks.

Scope: A single method paper presenting LATTLE (LLM Attention Transplant for Transfer LEarning), evaluated on ten source–target tabular data set pairs against 12 baselines.

Note on completeness: The provided paper content is truncated partway through Section 5.2 (at Figure 2). The remaining results in Section 5, the limitations discussion in Section 6, and the conclusion in Section 7 are not available, so those material are reported here as not reported.

What This Paper Is About

Tabular data sets from different domains rarely share the same columns, so transferring knowledge between them is hard — unlike images or text, which have homogeneous structure. The authors ask whether the internal attention weights of a language model, once adapted to a source table, can be physically transplanted into a tabular transformer to help classify a target table that shares no features with the source. Their goal is cross-domain transfer learning without shared features, without manual prompt engineering, and without enormous pretrained models.

Key Contributions

  1. A cross-attention transfer learning framework between an LLM and a feature-tokenized tabular transformer. The authors describe it as one of the first cross-attention transfer learning methods between an LLM and a feature tokenized transformer for tabular data.
  2. A tabular-specific training objective for the LLM. The language model is trained with a supervised cross-entropy classification loss designed for tabular data, rather than the standard auto-regressive next-token prediction loss.
  3. Attention-weight transplant instead of representation alignment. Context between heterogeneous domains is captured by transplanting the key and value projection weights of the LLM's topmost attention layer into the base layer of a downstream gated Feature Tokenizer Transformer (gFTT), which removes the requirement for shared features between source and target tables.
  4. Downstream learning without prompt engineering or in-context learning, demonstrated with a lightweight LLM (DistilGPT2) and a moderately sized source data set, rather than millions of samples and heavy compute.

Main Findings

  • LATTLE ranks first by average rank on AUC. Across the ten source–target pairs, LATTLE achieves an average rank of 2.0 (0.9), the best among all methods; the next best is TransTab at 4.8 (2.5), followed by XTab at 4.9 (3.3), CM2 at 5.2 (1.7), LLM (Single-domain) at 5.3 (3.8), and Logistic Regression at 5.4 (3.0).
  • LATTLE also ranks first by average rank on accuracy. Average accuracy rank is 2.3 (1.6) for LATTLE, versus TransTab 4.7 (3.0), LLM (Single-domain) 5.2 (4.1), XGBoost 6.1 (4.1), and XTab 6.3 (2.3).
  • The large pretrained tabular LLM performs worst on AUC. TABULA-8B, described as fine-tuned on eight billion tabular rows and evaluated at 32-shot (the maximum possible) in-context learning, has an average AUC rank of 10.4 (4.5), the lowest of the ranked methods; its accuracy scores are not reported (N/R).
  • Deep tabular architectures that lack transfer learning lag. ResNet averages rank 10.0 (3.3) on AUC and 8.3 (2.9) on accuracy; FT-Transformer ranks 8.9 (1.7) on AUC and 7.4 (4.1) on accuracy; TabNet ranks 9.9 (2.8) on AUC and 8.8 (3.4) on accuracy.
  • Transplanting only the topmost LLM layer into the lowest gFTT layer is the best strategy tested. In a controlled comparison (Table 4), the proposed one-layer transplant beat two alternatives: CG → DB scored 0.829 (0.04) versus 0.809 (0.05) for the W^L → W^{0,1} variant and 0.804 (0.06) for the two-layer variant; DG → VH scored 0.944 (0.00) versus 0.932 (0.00) and 0.936 (0.01); SK → CM scored 0.736 (0.03) versus 0.713 (0.03) and 0.719 (0.03).
  • Source data sets differ in how much transferable context they provide. Figure 3 is reported to show the effect of individual source data sets used in LLM pretraining on downstream transfer learning AUC, with cmc as the target data set; the numeric values of that figure are not reported in the available text.
  • In-context learning baselines were excluded. P2T is described as requiring extensive manual prompt engineering for sample-level source–target relationships, and in-context learning methods were excluded because the source code was unavailable. The paper also cites prior work showing LLM performance degrades when few-shot examples exceed 32.

Methodology in Plain English

The method is a two-stage pipeline.

Stage 1 – Teaching a language model about tables. A lightweight open-source LLM, DistilGPT2, is paired with a gated Feature Tokenizer Transformer (gFTT) and the two are fine-tuned together on the source table using a supervised cross-entropy classification loss (not the usual next-token prediction loss). Feature-value pairs are serialized into short text sentences, tokenized, and encoded through DistilGPT2's Byte Pair Encoding tokenizer. During this stage, the query, key, and value projection weights of the LLM's topmost attention layer and all five gFTT layers are updated. DistilGPT2 is chosen because it is open-source, cuts computational cost by 50% while retaining roughly 97% of GPT2's original performance, has six attention layers, and accepts up to 1024 input tokens.

Stage 2 – Transplanting attention into a new tabular model. The fine-tuned key and value projection weights from the LLM's topmost layer are copied into the base (lowest) attention layer of a new five-layer gFTT and frozen there. Only the query projection weight in that base layer stays trainable, so the target table's data "queries" the source domain's frozen key/value context. This is where cross-domain attention originates: attention between the source-adapted LLM and the target is computed at the level of projection weights rather than from raw query/key/value representations. The remaining gFTT layers are fine-tuned on the target table, and a [CLS] token's context vector feeds a linear classifier head whose logits are optimized with cross-entropy.

Evaluation setup. Fourteen public OpenML data sets were used: nine source and five target, spanning domains such as healthcare, finance, manufacturing, and software testing, with sample sizes from 540 to 70,000 and 8 to 76 features. A strict disjoint-feature-space constraint was enforced so no source and target share features — a stricter condition than prior transfer learning work. Ten source–target pairs were formed: CG-DB, CD-DB, MF-VH, DG-VH, CH-CM, SK-CM, SP-PC1, CE-PC1, SP-CB, and SB-CB. Each data set was split 70% training, 10% validation, 20% testing, sampled ten times with predefined random seeds; results are mean AUC and mean accuracy (ACC). DistilGPT2 was trained with a learning rate of 3e-4, weight decay 0.01, warm-up ratio 0.1, batch size 16, and 200 epochs on the source fold. The gFTT uses eight attention heads per layer and a feedforward network with hidden dimension 2048, ReLU activations, and dropout; hyperparameters were tuned with Optuna. Experiments ran on Ubuntu 20.04 with an Intel Core i9-13900F (32 cores, 5.60 GHz), 64 GB RAM, and an NVIDIA RTX 4090 GPU with 24 GB memory, using PyTorch.

Baselines. Twelve baselines: logistic regression, MLP, XGBoost, ResNet (adapted for tabular data), TabNet, FT-Transformer, TransTab, XTab, CM2, Tabula-8B, and two DistilGPT2 variants — one fine-tuned only on the target (LLM single-domain) and one first fine-tuned on source then adapted to target (LLM cross-domain). Traditional ML and non-transfer deep models were evaluated directly on target data with no source pretraining.

Why This Matters

Impact on research. The paper challenges two assumptions in tabular transfer learning: that source and target tables must share features, and that LLM-based tabular learning must go through prompt engineering or in-context learning. By transplanting attention weights rather than representations or prompts, it offers a mechanism-level alternative that appears to outperform much larger pretrained tabular models (Tabula-8B, trained on eight billion tabular rows) and dedicated cross-table pretraining frameworks (CM2, XTab) on average rank. It also suggests the appropriate learning objective for adapting a language model to non-text tabular data is a classification loss, not autoregressive next-token prediction.

Real-world applications (drawn from the domains represented in Table 1):

  • Healthcare: predicting metabolic disease (diabetes), heart disease (cardiovascular-disease), and thyroid disease (sick) when only limited labeled patient records are available.
  • Banking, telecommunications, and marketing: credit risk (credit-g), customer churn (churn), and demographics-based prediction (cmc).
  • Manufacturing and hazard monitoring: detecting steel plate faults (steel-plates-fault) and seismic bumps (seismic-bumps) for industrial safety.
  • Software testing and automotive/industrial design: defect-related tasks (pc1) and vehicle classification (vehicle), plus cylinder band design (cylinder-bands).

Industry relevance. Organizations hold many small, siloed tables with incompatible schemas. LATTLE requires no shared columns, no prompt writing, and no multi-billion-parameter model or massive pretraining corpus — only a lightweight open-source LLM (DistilGPT2) and a moderately sized source table. That profile is more practical for resource-constrained settings than fine-tuning or prompting a very large tabular language model.

Future Directions

  1. Test whether larger or newer open-source LLMs improve the transplant. The study deliberately uses DistilGPT2 (six attention layers); the authors note that more advanced open-source LLMs exist but may be memorization-prone or computationally expensive, leaving their transfer benefit unexplored here.
  2. Revisit the depth and placement of transplanted weights. Only the topmost LLM layer's key/value weights were transplanted into the base gFTT layer, and this beat the two tested alternatives in Table 4; whether

Authors’ abstract

Transfer learning on tabular data is challenging due to disparate feature spaces across domains, in contrast to the homogeneous structures of image and text. Large language models (LLMs) offer a knowledge base to improve the limited effectiveness of cross-domain transfer learning for tabular data. However, LLM performance often stagnates due to subjective text prompts and the computational limitations of in-context learning. We present a novel language-to-tabular context-learning method that uses attention-specific transformer weights, enabling seamless transfer learning across disparate tabular data sets. The LLM attention transplant mechanism facilitates a domain-agnostic transfer learning, eliminating the need for shared features between tables, LLM prompt engineering, and large-scale pretrained models. Our experiments using ten pairs of disjoint source-target data sets and 12 baseline methods demonstrate the superiority of the proposed LLM-attention transplant for transfer learning (LATTLE) method over traditional ML models, state-of-the-art deep tabular architectures, and models trained on thousands to billions of tabular samples. The proposed cross-domain attention transfer demonstrates an effective solution for adapting LLMs to learning non-text tabular data in a low-resource environment. The source code of the LATTLE implementation is publicly available.

Read the original paper