Skip to content
AI.info

Research

vLinear: A Powerful Linear Model for Multivariate Time Series Forecasting

vLinear: A Powerful Linear Model for Multivariate Time Series Forecasting Overview Research area: Machine learning / time series analysis — specifically multivariate time series forecasting using effi

vLinear: A Powerful Linear Model for Multivariate Time Series Forecasting
arXiv
2601.13768
Published
2026-01-20
Authors
Wenzhen Yue, Ruohao Guo, Ji Shi, Zihan Hao, Shiyu Hu, Xianghua Ying

AI summary

vLinear: A Powerful Linear Model for Multivariate Time Series Forecasting

Overview

  • Research area: Machine learning / time series analysis — specifically multivariate time series forecasting using efficient linear architectures combined with flow matching objectives.
  • Technical level: Intermediate to Advanced. The core ideas (a learnable vector replacing an attention matrix; a reweighted flow matching loss) are conceptually simple, but the paper grounds them in rank-1 matrix theory, maximum-likelihood derivations, and asymptotic complexity analysis.
  • Scope: The paper proposes a two-part forecaster — an O(N) token-dependency module called vecTrans and a final-series-oriented flow matching objective called WFMLoss — and evaluates it across 22 datasets and 124 forecasting settings.

Authors: Wenzhen Yue, Ruohao Guo, Ji Shi, Zihan Hao, Shiyu Hu, and Xianghua Ying (corresponding author), all at the State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University. arXiv:2601.13768v1 [cs.LG], 20 Jan 2026.

What This Paper Is About

State-of-the-art time series forecasters usually rely on self-attention (or close variants) to capture relationships between different variables, but this costs O(N²) computation and memory in the number of variates N. Purely linear models are far cheaper but tend to be less expressive, particularly at modeling cross-variate structure. This paper's goal is to close that gap: build a model that keeps the efficiency of a linear architecture while matching or beating attention-based forecasters in accuracy.

Key Contributions

  1. vecTrans — a linear-complexity token dependency module. Instead of an input-dependent attention matrix, vecTrans uses a single learnable vector of length N to aggregate token features and broadcasts the result back to every token. The paper proves (Theorem 1) that a row-wise L1-normalized rank-1 matrix must take the form 1aᵀ, which motivates this design, and by rearranging the order of operations it avoids ever materializing the N×N matrix — reducing complexity from O(N²) to O(N). It is a plug-in replacement for attention in Transformer-based forecasters.

  2. WFMLoss — a final-series-oriented, weighted flow matching objective. Typical flow matching objectives are "velocity-oriented": they supervise the instantaneous direction of the flow field, which depends on the initial noise sample. The authors instead supervise the final predicted series Ŷ⁽¹⁾ directly, and add path-weighting ((2−t)^(−0.5)) and horizon-weighting (i^(−0.5)) so learning focuses on the more reliable stages of the generation path and the near-future horizons. Theorem 2 justifies the horizon weighting as optimal under a maximum-likelihood criterion assuming a Laplace conditional density.

  3. A simplified, deterministic inference rule. Theorem 3 shows that if the velocity network is linear in Ŷ⁽ᵗ⁾ (as WFMLin is), running inference from the all-zero state is equivalent to computing the expectation over Gaussian noise samples — so the model needs no noise sampling at test time. This is verified empirically in Table 3, where all-zero, size-10, size-30, and size-50 Gaussian sampling give near-identical average MSEs.

  4. Large-scale empirical validation. 22 real-world datasets and 124 forecasting settings, with additional plug-in experiments showing both vecTrans and WFMLoss improve existing forecasters.

Main Findings

  • vecTrans beats attention and its variants. In Table 8, vecTrans achieves the best average MSE on ETTh1 (0.416), ETTm2 (0.268), ECL (0.153), Solar (0.209), Weather (0.233), and PEMS03 (0.093) when compared against vanilla self-attention, Gated Attention, E.Attn., Flow., Flash., FLattn., L.Attn., and Mamba inside the same vLinear framework.

  • Rank-1 is enough. Relaxing the rank-1 constraint to rank k = 4, 16, or 32 yields similar or slightly worse results (Table 9), e.g., ETTh1 MSE is 0.416 at k = 1 versus 0.420, 0.419, and 0.419 at k = 4, 16, 32 respectively. The authors conclude a rank-1 matrix suffices to capture multivariate correlations.

  • Final-series-oriented flow matching beats velocity-oriented. Table 2 shows WFMLoss outperforms the path- and horizon-weighted velocity objective by 8.6% on average (e.g., ETTh1 MSE 0.416 versus 0.434). The plain velocity MSE variant is much worse still (ETTh1 0.439, Solar 0.963, Weather 3.257), which the authors attribute to training instability.

  • WFMLoss beats other state-of-the-art losses. In Table 10 it outperforms the second-best loss by 3.2% on average across ETTh1, ETTm2, ECL, Solar, Weather, and PEMS03.

  • Both weightings help. Removing path-weighting, horizon-weighting, or switching from MAE to MSE degrades performance by 6%, 1%, and 14% respectively (Table 11).

  • Representation learning ablation. vecTrans-MLP outperforms the self-attention counterpart (Attn.-MLP) by 3.4% on average, and removing vecTrans entirely causes an 11% performance degradation. On large-variate datasets (ECL, Solar, PEMS03), vecTrans leads the second-best variant by about 5%.

  • The learned vector recovers the attention structure. Figure 6 shows the learned vector approximates the column-wise average of attention and NormLin matrices on the Weather dataset.

  • Top-1 counts. In long-term forecasting (Table 4, Input-96-Predict-{96, 192, 336, 720}, with PEMS using H ∈ {12, 24, 48, 96}), vLinear records 8 first-place MSE counts and 8 first-place MAE counts across the listed baselines. In short-term forecasting (Table 5), it records 5 and 5.

  • Efficiency. vecTrans has FLOPs of O(2ND² + ND) and memory O(ND), versus O(2N²D + 4ND²) FLOPs and O(hN² + ND) memory for multi-head self-attention, and O(N²D + 2ND²) / O(N² + ND) for the NormLin module (Table 1). When plugged into Transformer-based forecasters such as iTransformer, vecTrans delivers up to 5× inference speedups with consistent accuracy gains. Figure 4 shows vLinear uses fewer FLOPs and less GPU memory at the Input-96-Predict-96 setting.

  • Generalizability. vecTrans improves iTransformer, PatchTST, Leddam, and Fredformer when swapped in for attention (Table 12), and WFMLoss improves OLinear, iTransformer, PatchTST, and DLinear (Table 13) — e.g., PatchTST on Traffic improves from 0.531 to 0.470 MSE with WFMLoss.

  • Competitive on small data and strong baselines. vLinear also outperforms additional baselines SimpleTM, TQNet, TimePro, and TimeBase on most datasets (Table 6), and shows better robustness to random seeds than TimeMixer++ and iTransformer (Appendix G, as reported).

Methodology in Plain English

The model is deliberately simple on the inside. An input series is instance-normalized to counter non-stationarity, then passed through an orthogonal transform (OrthoTrans, borrowed from the prior OLinear work) that decorrelates the channels. A dimension-extension step expands the series into a richer representation, which is flattened and embedded.

The heart of the encoder is a stack of blocks that alternate between two cheap operations: (1) vecTrans, which handles relationships across variables, and (2) a two-layer MLP, which handles patterns across time. vecTrans works by taking one learnable vector of length N, passing it through a sigmoid and an L1 normalization so it behaves like a proper averaging weight, using it to compute a weighted summary of all variable tokens, and then broadcasting that summary back to every variable. Because the code multiplies this vector into the feature matrix before any N×N matrix would be formed, the cost scales linearly rather than quadratically.

After the encoder produces a conditional representation of the history, a small "flow block" generates the forecast. Rather than predicting the answer in one shot, it models a straight-line path from noise to the target series. At any interpolation time t, a single linear layer takes the history condition, the current intermediate state, and t, and outputs a velocity. The key trick is that the loss does not compare that velocity to the "true" velocity as in standard flow matching; instead it pushes the state all the way to t = 1 and compares that final prediction against the ground truth, using a weighted mean absolute error where later path points and nearer horizons get more weight. At test time, because the velocity network is linear, the whole sampling loop collapses into a deterministic-weighted sum of linear evaluations starting from an all-zero state — no random noise needed.

The authors train with PyTorch and the ADAM optimizer, and compare against 13 established baselines spanning linear-based models (OLinear, TimeMixer, FilterNet, FITS, DLinear), Transformer-based models (TimeMixer++, Leddam, Fredformer, iTransformer, PatchTST, CARD), and a TCN-based model (TimesNet).

Why This Matters

Impact on research. The paper challenges a working assumption in the field — that capturing multivariate correlations requires the quadratic cost of attention. If a single learnable vector can match or beat attention-based modules, then a large body of work on efficient attention variants is worth re-examining. It also reframes flow matching for forecasting: the paper's argument that velocity-oriented supervision forces the network to "memorize the starting noise" is a general critique that applies beyond this specific model, and WFMLoss is positioned as a drop-in objective that improves OLinear, iTransformer, PatchTST, and DLinear.

Real-world applications (drawn from the paper's own framing and datasets):

  • Weather forecasting — the Weather benchmark is one of the primary evaluation sets; improved long-horizon accuracy at lower compute directly affects operational forecasting costs.
  • Energy systems — the ECL (electricity consumption), Solar-Energy, and Power datasets represent grid load and generation forecasting, where many variables (regions, meters, plants) must be modeled jointly.
  • Transportation — the PEMS03/04/07/08 and METR-LA datasets cover traffic flow forecasting, a classic high-variate setting where linear scaling in the number of sensors matters.
  • Finance and economics — the Exchange, NASDAQ, SP500, DowJones, and Wiki datasets represent financial

Authors’ abstract

In this paper, we present \textbf{vLinear}, an effective yet efficient \textbf{linear}-based multivariate time series forecaster featuring two components: the \textbf{v}ecTrans module and the WFMLoss objective. Many state-of-the-art forecasters rely on self-attention or its variants to capture multivariate correlations, typically incurring $\mathcal{O}(N^2)$ computational complexity with respect to the number of variates $N$. To address this, we propose vecTrans, a lightweight module that utilizes a learnable vector to model multivariate correlations, reducing the complexity to $\mathcal{O}(N)$. Notably, vecTrans can be seamlessly integrated into Transformer-based forecasters, delivering up to 5$\times$ inference speedups and consistent performance gains. Furthermore, we introduce WFMLoss (Weighted Flow Matching Loss) as the objective. In contrast to typical \textbf{velocity-oriented} flow matching objectives, we demonstrate that a \textbf{final-series-oriented} formulation yields significantly superior forecasting accuracy. WFMLoss also incorporates path- and horizon-weighted strategies to focus learning on more reliable paths and horizons. Empirically, vLinear achieves state-of-the-art performance across 22 benchmarks and 124 forecasting settings. Moreover, WFMLoss serves as an effective plug-and-play objective, consistently improving existing forecasters. The code is available at https://anonymous.4open.science/r/vLinear.

Read the original paper