Research
OLR-WA: Online Weighted Average Linear Regression in Multivariate Data Streams
Overview Research area: Online / streaming machine learning, specifically multivariate online linear regression. Technical level: Intermediate. The method itself is conceptually simple (a weighted ave

- arXiv
- 2512.14892
- Published
- 2025-12-16
- Authors
- Mohammad Abu-Shaira, Alejandro Rodriguez, Greg Speegle, Victor Sheng, Ishfaq Ahmad
AI summary
Overview
Research area: Online / streaming machine learning, specifically multivariate online linear regression.
Technical level: Intermediate. The method itself is conceptually simple (a weighted average of two regression models), but the paper's geometric derivation uses N-dimensional hyperplanes, plane intersection cases, and pseudo-inverse linear algebra, and the evaluation compares against several classic online learners.
Scope: The paper introduces and empirically evaluates OLR-WA (OnLine Regression with Weighted Average), an N-dimensional online linear regression model that merges an incrementally built model with a retained base model, and benchmarks it against batch regression and seven established online regression algorithms on 14 datasets, including two adversarial scenario categories.
What This Paper Is About
Batch linear regression needs all the data available at once, has no time constraints, and assumes the data distribution never changes, which makes it a poor fit for streaming or evolving data. The authors extend their earlier OLR-WA work (which handled only 2-D and 3-D cases) to N-dimensional regression, building a model that updates incrementally from small mini-batches without storing the full history. The goal is to show that this weighted-average approach matches batch regression accuracy, beats or matches other online regressors, converges quickly even from a tiny initial sample, and is the only model tested that can stay deliberately conservative when older, higher-confidence data should be trusted over new data.
Key Contributions
-
A multivariate (N-dimensional) formulation of OLR-WA. The model combines a base model built from an initial batch with an incremental model built from each new mini-batch, merging them through a dynamically weighted average controlled by user-set weights W_base and W_inc.
-
A geometric construction for merging two hyperplanes. The paper defines two weighted average normal vectors (V_avg1 and V_avg2), corresponding to the two "sides" of the intersection of the base and incremental planes, generates candidate hyperplanes, and selects the better one by mean squared error over a sample combined from the incremental data and data sampled from the existing base model.
-
Explicit treatment of the plane-relationship edge cases. The authors enumerate the non-parallel case, the "Coincide" case (the planes align, so the algorithm skips to the next mini-batch), and the "Parallel" case (no intersection, so a weighted midpoint between the two models is used). They note these two cases are exceedingly rare and were not encountered in their experiments.
-
Adversarial evaluation split into two categories. The paper defines Time-Based adversarial scenarios (data drifts, and the model should follow the new distribution) and Confidence-Based adversarial scenarios (drift is present, but the model should stay conservative and favor older, trusted data), and reports that OLR-WA is the only tested model that performs well on the confidence-based category.
Main Findings
-
Batch regression was most accurate on the normal-regression datasets. In Table I (DS1 through CCPP), the batch pseudo-inverse model achieved the highest precision on every dataset.
-
OLR-WA tracked batch regression very closely. The paper reports the gap is very slight and almost in the third digit after the decimal point across all datasets, with one exception: on the 1KC dataset the batch model scored 0.93615 while OLR-WA scored 0.90773.
-
Several online baselines degraded sharply on specific datasets. RLS reached approximately 0.62667 on DS4 and 0.66197 on CCPP. PA reached approximately 0.78565 on 1KC against a batch score of 0.93615, and 0.66327 on CCPP against a batch score of 0.92855. Widrow-Hoff (LMS) reached 0.60772 on 1KC.
-
In time-based adversarial scenarios (DS5, DS6), OLR-WA, LMS, PA, and RLS adapted well. OLR-WA scored 0.98498 on DS5 (LMS 0.98528, RLS 0.98546, PA 0.97812) and 0.93634 on DS6 (LMS 0.93810, PA 0.91145, RLS 0.87134). SGD, MBGD, ORR, OLR, and the standard batch model produced minus R-squared (marked N/A) in this scenario.
-
In confidence-based adversarial scenarios (DS7, DS8), only OLR-WA produced a result. OLR-WA scored 0.97815 on DS7 and 0.93191 on DS8; every other model, including the batch model, is marked N/A.
-
OLR-WA converges from very few data points. On DS9 (Table III), OLR-WA scored 0.86265 using only 10 training points, versus SGD 0.32402, MBGD 0.03395, LMS 0.25216, ORR 0.28438, OLR 0.27546, RLS 0.13441, and PA 0.76154. OLR-WA's final value was 0.91293.
-
The base model size (BK) has little effect on performance. The paper reports robust performance with as few as 1% to 10% of the total training data points used for the base model.
-
Mini-batch size guidance. Because the method uses the pseudo-inverse, K must be at least M (the number of dimensions); the authors recommend a multiplier of at least 4, which yields results nearly identical to standard batch regression.
-
Computational complexity. Batch pseudo-inverse regression is O(NM²) for N samples and M features, whereas OLR-WA is approximately O(KM²) per iteration, where K is the mini-batch size and stays small and fixed as data accumulates.
-
Weight presets map to scenarios. Equal weights (W_base = 0.5, W_inc = 0.5) work as a default and are argued to guarantee convergence via the Exponentially Weighted Moving Average interpretation; dynamic weights can scale with the number of accumulated versus incremental points; a time-based setting such as W_inc = 2 and W_base = 0.1 gives new data 20 times the weight; a confidence-based setting such as W_inc = 0.1 and W_base = 2 gives old data 20 times the weight.
Methodology in Plain English
The authors build the model in three stages.
First, an initial "base" linear regression is fitted on a starting batch of data using the pseudo-inverse, which avoids explicitly inverting a matrix. Then, as data streams in, each new mini-batch is fitted into a separate "incremental" regression model. Crucially, the old data is never stored; instead, only the base model's coefficients are kept, and new data is evaluated against a sample generated from that base model.
Second, the base and incremental models are merged. Each linear regression is a hyperplane, characterized by a normal vector and a point. The authors compute two candidate normal vectors by taking weighted averages of the base and incremental normal vectors, using one sign flip to represent the two sides of where the planes intersect. They find a point that lies on both planes, then use that point plus each candidate normal vector to define a new hyperplane. If the planes are parallel instead of intersecting, they fall back to a weighted midpoint between the two models.
Third, the algorithm picks the better of the two candidate hyperplanes by comparing their mean squared error on a combined evaluation set, made from the new incremental data plus data sampled from the retained base model. The winner becomes the new base model for the next iteration.
The evaluation protocol used 5 random seeds for seed averaging and 5-fold cross-validation, so each reported r-squared value was validated across 25 experiments. All model weights were initialized to an array of zeros, feature engineering was kept minimal (normalization and one-hot encoding), and hyper-parameters were tuned per model and per dataset. Four hyper-parameters control OLR-WA: W_base, W_inc, K (mini-batch size), and BK (base model size), with K constrained by K = max(U, (M × Z+)) where U is the user-defined batch size and Z+ is a natural number of at least 1.
Why This Matters
Impact on research. The paper argues that the confidence-based adversarial scenario is a gap in existing online regression evaluation, and claims OLR-WA is the only model tested that handles it. It also shows that an online model can match batch regression accuracy on 14 datasets while running at O(KM²) per iteration instead of reprocessing all accumulated data.
Real-world applications (as described in the paper):
- Machine translation, where both a previously trained model and newly arriving training data carry importance (the paper's example for dynamic weight assignment).
- Sentiment analysis, where some labeled data points have been verified by experts or trusted sources and can therefore be weighted more heavily.
- Retail inventory modeling, illustrated by an Amazon-style store with an extensive product pool and hundreds of new products arriving daily, where a confidence-based weighting favors the larger established pool.
- Any time-based streaming setting where recent data should dominate, such as following a drifting data distribution.
Industry relevance. The memory savings matter for production systems: the method never needs all previously seen points resident, and the mini-batch size stays fixed as the stream grows, so per-iteration cost does not scale with accumulated history. The tunable weights also give practitioners a single knob to trade responsiveness to drift against trust in an established model.
Future Directions
-
Theoretical convergence guarantees. The paper states that the default equal-weight setting "guarantees" convergence via the Exponentially Weighted Moving Average, but the provided content does not report formal convergence-rate bounds. Establishing these, and comparing them with RLS and PA, is a natural next step.
-
Testing on truly high-dimensional and larger-scale data. The experiments span low to high dimensions across 14 datasets, but the truncated content does not report scaling tests to very large feature counts or very long streams, where the M × N pseudo-inverse cost may bind.
-
Handling the Coincide and Parallel cases in practice. The authors note these cases are exceedingly rare and were never encountered in their experiments, so their fallback behaviors (skip, or weighted midpoint) remain empirically untested.
-
Extending beyond linear regression. The weighted-average merging idea is described in purely geometric terms for hyperplanes; whether it transfers to nonlinear or kernelized models is not addressed in the provided content.
Target Audience
Researchers and practitioners working on streaming or online learning, incremental regression, and concept drift; engineers building real-time prediction systems on continuously arriving data where storing the full history is impractical; and readers interested in the specific problem of weighting trusted historical data against freshly observed data. Readers should be comfortable with linear algebra at the level of hyperplanes, normal vectors, and pseudo-inverses, and with standard regression evaluation metrics such as
Authors’ abstract
Online learning updates models incrementally with new data, avoiding large storage requirements and costly model recalculations. In this paper, we introduce "OLR-WA; OnLine Regression with Weighted Average", a novel and versatile multivariate online linear regression model. We also investigate scenarios involving drift, where the underlying patterns in the data evolve over time, conduct convergence analysis, and compare our approach with existing online regression models. The results of OLR-WA demonstrate its ability to achieve performance comparable to the batch regression, while also showcasing comparable or superior performance when compared with other state-of-the-art online models, thus establishing its effectiveness. Moreover, OLR-WA exhibits exceptional performance in terms of rapid convergence, surpassing other online models with consistently achieving high r2 values as a performance measure from the first iteration to the last iteration, even when initialized with minimal amount of data points, as little as 1% to 10% of the total data points. In addition to its ability to handle time-based (temporal drift) scenarios, remarkably, OLR-WA stands out as the only model capable of effectively managing confidence-based challenging scenarios. It achieves this by adopting a conservative approach in its updates, giving priority to older data points with higher confidence levels. In summary, OLR-WA's performance further solidifies its versatility and utility across different contexts, making it a valuable solution for online linear regression tasks.