Research
From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification
From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification Overview Research area: Machine learning — interpretable tabular data classification, knowledge distilla

- arXiv
- 2610.10227
- Published
- 2026-10-07
- Authors
- Yue Qiu, Zekang Du, Yiqun Diao, Bingsheng He, Qinbin Li
AI summary
From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular ClassificationOverview
Research area: Machine learning — interpretable tabular data classification, knowledge distillation from large language models (LLMs) into decision trees, few-shot learning.
Technical level: Intermediate. The core pipeline is conceptually approachable, but the paper includes a PAC-style generalization bound, Gini-impurity split criteria, and a formal problem statement.
Scope: The paper proposes LLMT, a three-stage framework that distills LLM knowledge into a standalone decision tree for few-shot tabular classification, and evaluates it on 11 public datasets against 17 baselines.
What This Paper Is About
LLMs hold broad world knowledge but are costly and opaque when used directly to classify tabular data, while decision trees are fast and interpretable but need substantial labeled data. The paper asks whether LLM knowledge can be re-expressed as a decision tree in few-shot settings, so the resulting model inherits the LLM's prior while keeping tree interpretability and cheap inference. Instead of prompting an LLM to emit a whole tree or a long reasoning path at once — which the authors find unstable and expensive — LLMT prompts for discrete rules and organizes them statistically.
Key Contributions
- A three-stage distillation paradigm that bridges the gap between expensive, opaque LLM reasoning and efficient, interpretable tree models, comprising Rule Generation, Tree Assembly, and Leaf Refinement.
- A formal generalization analysis of LLMT under explicit structural assumptions, presented as a PAC-style generalization error bound (Theorem 4.2), clarifying the method's statistical properties.
- Extensive experiments on 11 datasets with 17 baselines, reporting accuracy and efficiency across different few-shot settings, including high-dimensional datasets and an optional ensemble extension (LLMT Forest).
- Demonstrated practical gains: consistent accuracy improvements, large token-cost savings, and interpretable outputs in the few-shot setting, with code released at github.com/yueqiu0/LLMTree.
Main Findings
- Rule generation beats path generation: Prompting LLMs to produce discrete rules yields higher Set Utility and Set Accuracy than generating reasoning-path sets with the same total number of atomic feature conditions, across eight datasets. The authors conclude LLMs struggle to produce coherent hierarchical reasoning paths in few-shot settings.
- LLMs rely on metadata: On Diabetes and Spambase under zero-shot prompting, prediction accuracy of four LLMs substantially deteriorates when feature names and task descriptions are masked; the comparison is extended to six newer datasets in the appendix.
- Best average accuracy at two shots per class: In Table 2, LLMT reaches an average of 0.753 across eight datasets, versus 0.701 for the strongest baseline FeatLLM, 0.676 for LogReg, 0.675 for SVM, 0.637 for CART, 0.634 for CoT-Tree, 0.626 for GPTree, 0.626 for DirectZSDT, 0.623 for IO-Tree, 0.617 for DeLTa, 0.612 for LightGBM, 0.609 for ToT-Tree, 0.587 for XGBoost, and 0.564 for StepZSDT. LLMT leads on Nursery (0.762), Diabetes (0.764), and Spambase (0.811).
- Reported margins: The paper states LLMT outperforms almost all baselines, with more than 10% improvement over the strongest baseline in many settings.
- Strong in the extreme few-shot regime: At 3 shots on Nursery, LLMT scores 0.762 against 0.539 for FeatLLM; at 2 shots on Diabetes it scores 0.662; at 32 shots on Spambase it reaches 0.843 versus 0.837 for FeatLLM and 0.835 for LogReg.
- Token and time savings versus ToT-Tree: At 32 shots, LLMT uses 3.3k total tokens (Diabetes), 3.6k (Nursery), and 6.4k (Spambase), corresponding to 10.1x, 10x, and 22.1x token savings and 1.63x, 2.03x, and 3.56x construction-time savings relative to ToT-Tree.
- Cheaper than several LLM-assisted baselines: LLMT trains in 20.48 s (Diabetes), 16.33 s (Nursery), and 18.18 s (Spambase); FeatLLM takes 259.07 s, 416.16 s, and 324.06 s, and GPTree takes 140.75 s, 103.47 s, and 125.56 s. StepZSDT is far more expensive at 1006.67 s to 4036.18 s.
- Most concise, purest trees: LLMT achieves the lowest average Gini impurity of rules among compared methods (0.401 average), best on Diabetes (0.396) and Nursery (0.456), and near best on Spambase (0.352, versus 0.341 for IO-Tree). LLMT produces the most concise trees on Diabetes and remains among the most compact on Nursery and Spambase.
- Leaf refinement helps: Removing Leaf Refinement reduces accuracy by 8% on Diabetes and 2% on Spambase. LLMT is generally stable across different threshold values of tau, with tau = 0.7 usually best.
- Recursive prompting can hurt: ToT-Tree often performs worse than CoT-Tree, which the authors attribute to noise and inconsistency from recursive node-level prompting on structured tabular data.
- High-dimensional robustness: Results on three datasets with over 100 attributes show LLMT maintains its advantage in high-dimensional feature spaces, and support the ensemble extension.
Methodology in Plain English
LLMT distills an LLM into a decision tree through three stages.
- Rule Generation. The LLM is given a system prompt with task constraints and a dataset prompt containing feature names, feature types, and dataset-agnostic rule-format examples. It outputs discrete classification rules of the form
f_j op theta, where the operator is one of<,>=,=,!=and theta is a numerical threshold or categorical value; categorical features are handled by one-vs-rest splits. Each rule carries a self-assessed confidence score on a scale of 1 to 10. Crucially, only feature semantics — not raw sample values — are sent to the LLM, which the paper describes as preserving sample-level privacy. - Tree Assembly. The candidate rules are organized into a tree level by level using the small labeled set. At each node, rules already used on the path to that node are excluded, and remaining rules are grouped into bins by their confidence score, processed in descending order. Within a group, the rule that minimizes the weighted Gini impurity of the resulting split is selected, with ties broken by highest Gini gain and then lowest index. Nodes follow a breadth-first heap indexing scheme (root 1, children 2 and 3, and so on). Construction stops when the maximum depth is reached or no remaining rule yields positive gain. The intuition is to place important rules high in the tree so they are shared across more reasoning paths, while using few-shot data to adjust positions.
- Leaf Refinement. For every root-to-leaf path, the LLM is prompted with a natural-language description of the rule chain and returns a confidence distribution over labels. If the top confidence exceeds a threshold tau (0.7 in the experiments) and its label differs from the current leaf label, the leaf is changed. The paper notes this can also identify same-label sibling leaves for merging, simplifying the tree.
An optional LLMT Forest extension handles high-dimensional data: features are randomly permuted and partitioned into disjoint, approximately equal-sized subsets (quantity q = ceil(sqrt(|F|))), one LLMT tree is trained per subset on the same few-shot samples, and predictions are aggregated by majority vote.
Experimental settings: a Linux server with 4× Intel Xeon Gold 5117 CPUs and 4× Nvidia Tesla V100 GPUs; Qwen2.5-72B-Instruct via TogetherAI by default; maximum tree depth 3 or 4 depending on the dataset; number of rules K = max(10, 2^d − 1) where d is the maximum depth; leaf-refinement threshold tau = 0.7 for all datasets; 10 trials with randomly sampled training sets and a fixed test set of 100 samples, reporting mean ± standard deviation. The theoretical analysis includes a PAC-style bound showing that tree depth d is the dominant complexity driver through its 2^d dependence, which the authors use to justify shallow trees.
Why This Matters
Impact on research. The work argues for separating knowledge extraction from model construction: rather than making an LLM the final classifier or asking it to emit a whole tree, LLMT extracts rules and assembles them statistically. It contributes a finite, statistically analyzable hypothesis space — distilling the LLM into a finite rule set of size K — in contrast to black-box prompting pipelines, and provides a generalization bound that explicitly ties error to tree depth and rule-set size. It also positions itself against zero-shot tree builders (which cannot use labeled examples), data-hungry rule generators like GPTree, and LLM-at-inference methods like TabLLM, InsightTab, and SumBoost.
Real-world applications.
- Healthcare diagnostics, where interpretable decision paths support clinical justification and the paper's Glioma and Breast datasets are drawn from medical domains.
- Financial risk assessment, where decisions must be auditable and inference latency and cost matter.
- Customer behavior prediction, a domain the paper cites as a common tabular application.
- Resource-constrained or real-time deployment, where the paper argues LLM inference cost and latency are prohibitive and a compiled tree avoids calling the LLM at prediction time.
Industry relevance. The headline operational claim is cost: LLMT avoids repeated planning and voting, cutting token usage by 10–22x and construction time by 1.6–3.6x relative to ToT-Tree, while conducting no LLM calls at inference. The paper reports that LLMT preserves sample-level privacy by deriving rules solely from feature semantics without transmitting raw data values, which matters for regulated deployments. The stated limitation is that accuracy depends on the LLM having sufficient background knowledge of the target dataset, and the distilled tree inherits the LLM's biases — the authors advise fairness-aware prompting and post-hoc auditing for high-stakes use.
Future Directions
- Knowledge-scope dependence: The paper states that if a dataset lies outside the LLM's knowledge scope, the generated rules and the constructed tree may be inaccurate or ineffective. How to detect and compensate for this gap is left open.
- Bias and fairness: The method inherits biases in the underlying LLM, which may influence extracted rules and downstream decisions; the authors call for fairness-aware prompting and post-hoc auditing.
- Tightening the theory: The authors acknowledge the PAC-style bound can be numerically loose in the extreme few-shot regime where n is small, even though it gives structural guidance on depth.
- Extending the design space: The paper reports LLMT Forest for high-dimensional data and refers to appendix studies across varying tree depths and larger baseline sets, leaving room to explore ensemble construction, rule selection, and refinement strategies further.
Target Audience
Researchers and practitioners working on tabular machine learning, interpretable models, and LLM knowledge distillation; engineers who need accurate classification under few-shot constraints with low inference cost and auditability; and readers interested in formal generalization analysis of LLM-guided model construction. Familiarity with decision trees, Gini impurity, and few-shot prompting is helpful, and the appendix — containing proofs, algorithms, prompts, and full results — is aimed at those who want to reproduce or extend the method.
Authors’ abstract
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.