Research
Laplace Approximation For Tensor Train Kernel Machines In System Identification
Overview Research area: Machine learning for nonlinear system identification, at the intersection of Gaussian process regression, tensor-network (tensor train) models, and approximate Bayesian inferen

- arXiv
- 2512.02532
- Published
- 2025-12-02
- Authors
- Albert Saiapin, Kim Batselier
AI summary
Overview
- Research area: Machine learning for nonlinear system identification, at the intersection of Gaussian process regression, tensor-network (tensor train) models, and approximate Bayesian inference.
- Technical level: Advanced. The paper assumes familiarity with Gaussian processes, tensor decompositions, alternating least squares, Laplace approximation, and variational inference.
- Scope (one sentence): The paper develops LA-TTKM, a Bayesian tensor train kernel machine that places a Laplace approximation over one selected TT-core and uses variational inference for the precision hyperparameters, then evaluates core/rank/feature choices on seven UCI regression datasets and the inverse dynamics of a robotic arm.
What This Paper Is About
Gaussian process regression gives predictions with principled uncertainty, but inverting an N x N kernel matrix costs O(N³), which limits large-scale or real-time system identification. Tensor-network kernel machines avoid that cost by representing an exponentially large set of basis functions with a low-rank tensor train, but turning them into a fully probabilistic model raises unresolved design questions. This paper builds a Bayesian version, LA-TTKM, and answers those questions: which TT-core to treat as a random variable, whether the TT-rank pattern matters, and how to set the hyperparameters without cross-validation.
Key Contributions
- A Bayesian tensor train kernel machine (LA-TTKM). The method combines a Laplace approximation over a single selected TT-core with variational inference over the precision hyperparameters β and γ, yielding the probabilistic graphical model shown in Figure 1 and a closed-form Gaussian predictive distribution.
- Resolution of open questions from Menzen et al. (2023). The paper systematically investigates which TT-core should be treated in a Bayesian manner, how TT-ranks influence that choice, and whether input-feature ordering matters.
- Variational inference as a replacement for cross-validation. The paper reports that VI selects hyperparameters much faster than cross-validation (stated as up to 65x in the abstract and approximately 65x on average in the conclusion) while delivering comparable predictive quality.
- Demonstration on a large-scale inverse dynamics task. LA-TTKM is evaluated on an industrial robot dataset and compared against both a cross-validated version of itself and a full GP baseline.
Main Findings
- The Bayesian core choice is largely independent of the TT-rank pattern. Table 1 reports the best core out of the total number of cores for each dataset under three rank patterns. For example, on Boston the best core is 5/13 under Pattern 1, 7/13 under Pattern 2, and 9/13 under Pattern 3; on Kin8nm, Naval, Protein, and Energy the same core is best across patterns. Pattern 1 denotes ranks [R, R, R] with R > 1, Pattern 2 uses [R, P, R] with P > R, and Pattern 3 uses [R, P, R] with P < R.
- Feature ordering does not practically influence performance. Table 2 varies cyclic shifts of the input features (No Shift, Shift 1, Shift 2, Shift 3). The paper reports no universally optimal choice, and notes that the best-performing core tends to be either the first or the last TT-core because those cores contain the fewest parameters, since R₀ = R_D = 1.
- The first TT-core is adopted for scalability. Based on the two ablations, the paper recommends using a single uniform rank and the first TT-core as the Bayesian core, since it has the smallest number of parameters.
- VI trains much faster than cross-validation. Table 3 reports test NLL and training time over 10 independent restarts, with parallelization disabled. CV times range from 119.8 ± 8.5 s (Energy) to 1397.5 ± 63.0 s (Concrete); VI times range from 0.9 ± 0.4 s (Yacht) to 48.6 ± 11.5 s (Concrete). The abstract states up to 65x faster training, the conclusion states approximately 65x on average, while the Section 4.2 prose describes the average speed-up as approximately 65%.
- VI's predictive quality is dataset-dependent. In Table 3, VI gives lower NLL than CV on Concrete (0.26 ± 0.14 vs 0.38 ± 0.09), Kin8nm (0.38 ± 0.02 vs 0.51 ± 0.19), Naval (-1.16 ± 0.15 vs -0.13 ± 0.02), Protein (1.13 ± 0.01 vs 1.23 ± 0.00), and Yacht (-0.41 ± 0.09 vs 0.03 ± 0.22), and higher NLL on Boston (1.55 ± 0.83 vs 1.16 ± 0.05) and Energy (0.93 ± 0.62 vs -0.07 ± 0.01).
- On inverse dynamics, VI trades uncertainty calibration for accuracy and speed. On the robot dataset (39,988 training samples, 3636 test samples, 18-dimensional inputs, predicting a single torque component), LA-TTKM with VI achieves NLL 6.2, RMSE 1.1, and 22.8 s; with CV, NLL 2.6, RMSE 1.4, and 1364.2 s; and the full GP baseline achieves NLL 4.4, RMSE 1.8, and 7446.3 s. That is a 60x speedup over CV and 325x over full GP, with lower RMSE but higher NLL due to occasionally overconfident uncertainty estimates.
- Relationship to prior work. The paper notes that if one sets σ_y² = β⁻¹, Λ = γ⁻¹I, and ΦW = A⁽ᵈ⁾, the resulting predictive distribution coincides with the one derived in Menzen et al. (2023).
Methodology in Plain English
The model starts from a linear regression in a feature space, where the features are built as a tensor product of per-dimension feature maps. To keep this tractable, the weight vector is parameterized as a tensor train, so an exponentially large weight vector is stored with only a polynomial number of parameters. Training then alternates between cores using the ALS algorithm, which solves exactly for one core while the others are held fixed.
To make the model probabilistic, the paper treats only one TT-core as a random variable and keeps the rest deterministic. Because ALS finds the exact optimum for that one core, a Laplace approximation around that optimum is justified, giving a Gaussian posterior whose covariance is βAᵀA + γI and whose mean is exactly the ALS solution. The hyperparameters β and γ receive Gamma priors and are updated by coordinate-ascent variational inference rather than by a grid search, and their expectations are plugged back into the weight posterior. Because only a Gaussian appears inside the prediction integral, the predictive distribution over test outputs is available in closed form, with the design matrix for test inputs built the same way as during training.
Implementation notes: the paper uses unit-norm polynomial features or Fourier features, sets the number of basis functions uniformly across modes (I_d = I), chooses the Bayesian core to match the core at which ALS is terminated, and fixes all remaining hyperparameters so that the total parameter count stays below the number of training samples. The CV comparison searches γ and β over {0.001, 0.01, 0.1, 1, 10}. For reference, standard GP inference costs O(N³), the finite-basis-function approach costs O(NL²), and ALS training and memory are stated as T = O(E D I² R⁴ (N + I R²)) and M = O(N R (D + I R)), where E is the number of ALS epochs. Experiments ran on a Dell Latitude 7440 laptop with a 13th Gen Intel Core i7-1365U CPU and 16 GB RAM, except the full GP baseline, which ran on 2 x AMD EPYC 7252 8-core CPUs with 256 GB RAM. Code is available at https://github.com/AlbMLpy/laplace-ttkm.
Why This Matters
Impact on research. The paper connects two literatures that usually stay separate: tensor-network kernel machines and Bayesian GP approximations. It answers concrete design questions left open by Menzen et al. (2023), converts hyperparameter selection from a cross-validation grid search into a variational update, and shows that Laplace approximation remains theoretically defensible when only the core that ALS solves exactly is treated as random. The result that the Bayesian core and feature ordering barely matter removes a design decision that would otherwise require empirical tuning.
Real-world applications mentioned or implied by the paper's context:
- Inverse dynamics estimation for robotic arms, the task demonstrated on the Weigand et al. (2022) industrial robot dataset.
- Model predictive control, which the paper cites as a successful application area for GP regression (Hewing and Zeilinger, 2017).
- Nonlinear state estimation, cited through Särkkä (2019) and Bergmann et al. (2018).
- Large-scale or real-time nonlinear system identification, the scalability bottleneck the paper is explicitly written to address.
Industry relevance. Training times in the tens of seconds rather than minutes (Table 4: 22.8 s vs 1364.2 s vs 7446.3 s) matter for settings where models are retrained repeatedly. The tensor train parameterization also allows models whose nominal parameter count would be 8¹⁸ for the robotic arm experiment to be trained at all on a laptop-class machine when a full GP baseline is not feasible.
Future Directions
- Make more cores Bayesian. The paper proposes extending the framework to multiple, or all, TT-cores, or to a mean-field approximation, as richer probabilistic formulations.
- Automatic TT-rank adaptation. The authors suggest letting the model grow incrementally in rank as needed, without retraining from scratch.
- Calibrating VI uncertainty. The inverse dynamics results show VI achieving lower RMSE but higher NLL than CV and full GP, with narrow predictive intervals in Figure 2; improving uncertainty calibration while keeping the speed advantage is an open question raised by these numbers.
- Scaling and hyperparameter scope. The ablation experiments fixed all hyperparameters other than γ and β so that the parameter count stayed below the number of training samples, and the largest dataset used 39,988 training samples; how the design choices hold on larger or noisier data is not settled by the reported experiments.
Target Audience
This paper suits readers already comfortable with probabilistic machine learning and tensor methods: researchers working on Gaussian process approximations, tensor-network kernel machines, or Bayesian system identification; graduate students and practitioners who need uncertainty-aware models for control, robotics, or state estimation and who are willing to trade implementation complexity for scalability. Readers without a background in tensor trains, Laplace approximation, or variational inference will find the technical development in the main text difficult to follow, though the ablation tables and the inverse dynamics comparison are readable on their own.
Authors’ abstract
To address the scalability limitations of Gaussian process (GP) regression, several approximation techniques have been proposed. One such method is based on tensor networks, which utilizes an exponential number of basis functions without incurring exponential computational cost. However, extending this model to a fully probabilistic formulation introduces several design challenges. In particular, for tensor train (TT) models, it is unclear which TT-core should be treated in a Bayesian manner. We introduce a Bayesian tensor train kernel machine that applies Laplace approximation to estimate the posterior distribution over a selected TT-core and employs variational inference (VI) for precision hyperparameters. Experiments show that core selection is largely independent of TT-ranks and feature structure, and that VI replaces cross-validation while offering up to 65x faster training. The method's effectiveness is demonstrated on an inverse dynamics problem.