Skip to content
AI.info

Research

VaSST: Variational Inference for Symbolic Regression using Soft Symbolic Trees

Overview Research area: Statistical methodology (stat.ME) at the intersection of Bayesian inference, variational optimization, and symbolic regression for scientific machine learning. Technical level:

VaSST: Variational Inference for Symbolic Regression using Soft Symbolic Trees
arXiv
2602.23561
Published
2026-02-27
Authors
Somjit Roy, Pritam Dey, Bani K. Mallick

AI summary

Overview

Research area: Statistical methodology (stat.ME) at the intersection of Bayesian inference, variational optimization, and symbolic regression for scientific machine learning.

Technical level: Advanced. The paper assumes familiarity with Bayesian hierarchical modeling, conjugate Normal-Inverse-Gamma priors, variational inference and the ELBO, mean-field approximations, and continuous relaxations of discrete random variables (Binary-Concrete, Gumbel-Softmax).

Scope in one sentence: The paper proposes VaSST, a fully probabilistic symbolic regression framework that replaces discrete symbolic search with a differentiable "soft symbolic tree" representation optimized by black-box variational inference.

What This Paper Is About

Symbolic regression (SR) tries to recover concise closed-form mathematical expressions from data, but existing approaches are dominated by heuristic evolutionary search, data-hungry neural generators, or Markov chain Monte Carlo over discrete tree structures that mixes poorly in a combinatorially vast space. The authors introduce VaSST (Variational Inference for Symbolic Regression using Soft Symbolic Trees), which relaxes the discrete operator and feature assignments of a symbolic expression tree into probability distributions so that inference becomes a continuous, gradient-based optimization problem while remaining a coherent Bayesian model.

Key Contributions

  1. A variational inference framework for symbolic regression. VaSST recasts posterior inference over symbolic structures as optimization rather than Markov chain Monte Carlo, replacing a combinatorial search over an astronomically large discrete space with gradient-based exploration.

  2. Soft symbolic trees. Discrete operator and feature assignments in a full binary tree skeleton are replaced by continuous relaxations, using the Binary-Concrete relaxation for expansion indicators and Gumbel-Softmax for operator and feature assignments. This yields a differentiable surrogate for evaluating symbolic expressions through soft gating over unary and binary operators and leaf features.

  3. A depth-dependent regularizing tree prior and posterior-aware model selection. A Bernoulli split probability of the form p_ζ = α(1 + d_ζ)^{-δ} penalizes depth, encouraging parsimonious expressions. After optimization, symbolic structures are sampled from the learned soft representation, ranked by a posterior-aware evidence score, and summarized with an Occam's window-based posterior summary of high-support symbolic explanations.

  4. Empirical comparison against state-of-the-art SR methods. VaSST is evaluated on simulated experiments and the Feynman Symbolic Regression Database against competing methods drawn from the SRBench ecosystem, alongside a public Python implementation.

Main Findings

  • Structural recovery and predictive accuracy: On simulated experiments and the Feynman Symbolic Regression Database, VaSST is reported to achieve strong structural recovery and predictive accuracy compared to state-of-the-art competing SR methods. No numeric performance values, error metrics, or recovery rates appear in the available text.

  • Improved scalability over existing probabilistic SR: The authors argue that MCMC-based discrete structural exploration in earlier Bayesian SR methods limits scalability beyond small- to moderate-sized data settings. VaSST's continuous relaxation is presented as addressing this.

  • Diagnosed weaknesses in competing methods: Bayesian Machine Scientist (BMS) is described as using an ad hoc structural prior derived from corpus parsing based on a priori knowledge, with Metropolis-Hastings proposals over discrete tree structures that can mix poorly. Bayesian Symbolic Regression (BSR) is described as a partial Bayesian formulation using plug-in ordinary least squares estimates for outer regression coefficients, which the authors say incompletely propagates parameter uncertainty and frequently produces overly complicated output structures, as evidenced in the paper's Section 5.

  • Diagnosed weaknesses in machine-learning-based SR: Neural, search-driven SR methods remain computationally challenging due to the NP-hard nature of SR, and their effectiveness is described as often contingent on large training datasets and low-noise regimes, as demonstrated in Section 5. Large language model-guided SR methods are described as sensitive to prompt design and model choice and capable of producing syntactically invalid or unstable expressions.

  • Combinatorial scale of the discrete problem: The authors quantify the induced discrete model space as growing exponentially with tree depth D and ensemble size K, with O((2p|O|)^{N·K}) possible configurations, where N = 2^{D+1} − 1 is the number of skeleton nodes.

  • Parsimony mechanism: The depth-dependent split probability p_ζ = α(1 + d_ζ)^{-δ} is presented as the mechanism that imparts a regularizing effect on individual tree depths, in line with Occam's razor.

Methodology in Plain English

VaSST models the response as a linear combination of K symbolic expressions plus an intercept and Gaussian noise, so the fitted function is an ensemble of trees rather than a single expression. The regression coefficients and noise variance receive a conjugate Normal-Inverse-Gamma prior, which lets the authors integrate them out analytically and work with a marginal likelihood that depends only on the design matrix built from the symbolic expressions.

Each symbolic expression is embedded in a full binary tree skeleton of fixed maximum depth D, indexed in heap order, with a total of N = 2^{D+1} − 1 nodes. Every skeleton node carries three structural variables: an expansion indicator saying whether the node is internal or a leaf, an operator label if internal, and a feature label if a leaf. A deterministic pruning operation turns the skeleton into a valid symbolic tree by removing descendants of leaves and the right subtree of unary operators.

Priors are placed on these structural variables: a Bernoulli prior on expansion with depth-dependent probability, categorical priors on operators and features governed by Dirichlet-distributed weight vectors, and Dirichlet hyperpriors. To make inference tractable, the discrete structural variables are relaxed into continuous ones using Binary-Concrete and Gumbel-Softmax transformations with temperature parameters τ_ex, τ_op, and τ_ft. Node evaluation then becomes a smooth interpolation: a leaf contributes a convex combination of input features, an internal node mixes unary and binary operations over its children, and soft gating combines the three cases.

Because the expected log marginal likelihood has no closed form and the design matrix depends nonlinearly on the soft trees, the evidence lower bound (ELBO) is approximated by Monte Carlo sampling of soft structural variables, producing a differentiable objective. Variational parameters include sigmoid-transformed logits for expansion, softmax-transformed logits for operators and features, and Dirichlet parameters. The objective is maximized with gradient-based black-box optimization, with temperature annealing balancing exploration against sharpening toward discrete structures. The resulting variational distributions are then sampled to produce candidate symbolic expressions, which are ranked by a posterior-aware evidence score and summarized through an Occam's window. Algorithm boxes for node-level soft evaluation, full soft evaluation, and the approximate ELBO are provided in the appendices.

Why This Matters

Impact on research: The paper argues that fully probabilistic SR formulations remain scarce, and that existing Bayesian alternatives rely on discrete MCMC exploration that scales poorly. By turning symbolic structure learning into differentiable optimization, VaSST offers a route to uncertainty quantification over plausible symbolic forms, which prediction-centric regression models and heuristic search methods do not provide.

Real-world applications (drawn from the domains the paper cites as motivations for scientific machine learning and symbolic regression):

  • Materials science, where symbolic regression has been used to accelerate materials design.
  • Climate and weather prediction, where interpretable governing relationships matter for downstream decision-making.
  • Biology, where explicit mechanistic equations are preferred over black-box predictors.
  • Physics, including sparse discovery of nonlinear dynamical systems and recovery of fundamental scientific laws.

Industry relevance: Any setting where interpretable closed-form models are valued over opaque predictors stands to benefit, including engineering design, process modeling, and scientific software workflows. The availability of a Python implementation and evaluation against the SRBench ecosystem lowers the barrier for practitioners to benchmark VaSST against established baselines.

Future Directions

  • Temperature annealing strategy: The paper notes that careful annealing of τ_ex, τ_op, and τ_ft balances exploration and structural learning, but a principled schedule remains an open design question.
  • Approximation quality: The appendices discuss the approximation gap associated with the Monte Carlo estimate of the ELBO, leaving room for tighter bounds or lower-variance estimators.
  • Scaling behavior: The discrete configuration count O((2p|O|)^{N·K}) grows with feature count p, operator set size, depth D, and ensemble size K; how VaSST behaves at larger depths, larger ensembles, and higher-dimensional inputs is a natural extension.
  • Richer symbolic spaces and domain knowledge: The operator set used here is {sin, cos, exp, log, ^2, ^3} for unary and {+, ×, −, /} for binary operators, and the paper contrasts its approach with prompt-sensitive LLM-guided SR, suggesting possible integration of domain constraints or hybrid LLM priors. The specific accuracy, noise-robustness, and runtime figures underlying the empirical claims are not reported in the available text, and reproducing or extending the benchmark comparison would be a further step.

Target Audience

Statisticians and machine learning researchers working on Bayesian inference, variational methods, or symbolic regression will find the primary contribution in the soft symbolic tree formulation and its variational treatment. Practitioners in scientific machine learning who need interpretable equations with uncertainty estimates, and readers already familiar with SRBench baselines such as BMS, BSR, and VaRT, are the most likely to benefit from the empirical comparison and the accompanying Python implementation.

Authors’ abstract

Symbolic regression has recently gained traction in AI-driven scientific discovery, aiming to recover explicit closed-form expressions from data that reveal underlying physical laws. Despite recent advances, existing methods remain dominated by heuristic search algorithms or data-intensive approaches that assume low-noise regimes and lack principled uncertainty quantification. Fully probabilistic formulations are scarce, and existing Markov chain Monte Carlo-based Bayesian methods often struggle to efficiently explore the highly multimodal combinatorial space of symbolic expressions. We introduce VaSST, a scalable probabilistic framework for symbolic regression based on variational inference. VaSST employs a continuous relaxation of symbolic expression trees, termed soft symbolic trees, where discrete operator and feature assignments are replaced by soft distributions over allowable components. This relaxation transforms the combinatorial search over an astronomically large symbolic space into an efficient gradient-based optimization problem while preserving a coherent probabilistic interpretation. The learned soft representations induce posterior distributions over symbolic structures, enabling principled uncertainty quantification. Across simulated experiments and Feynman Symbolic Regression Database within SRBench, VaSST achieves superior performance in both structural recovery and predictive accuracy compared to state-of-the-art symbolic regression methods.

Read the original paper