Research
LOBERT: Generative AI Foundation Model for Limit Order Book Messages
Overview Research area: Machine learning for financial market microstructure — specifically generative, foundation-model-style sequence modeling of Limit Order Book (LOB) message streams (Level III da

- arXiv
- 2511.12563
- Published
- 2025-11-16
- Authors
- Eljas Linna, Kestutis Baltakys, Alexandros Iosifidis, Juho Kanniainen
AI summary
Overview
- Research area: Machine learning for financial market microstructure — specifically generative, foundation-model-style sequence modeling of Limit Order Book (LOB) message streams (Level III data), using an adaptation of the BERT architecture.
- Technical level: Advanced. The paper assumes familiarity with transformer architectures, tokenization, masked language modeling, rotary position embeddings, and order book terminology.
- Scope (one sentence): The paper introduces LOBERT, an encoder-only foundation model that treats complete LOB messages as single tokens and is pre-trained with Masked Message Modeling, then fine-tuned for next-message prediction and mid-price direction prediction on Nasdaq ITCH data for AAPL, INTC, MSFT and FB.
What This Paper Is About
Limit order book dynamics are driven by irregularly timed events — order submissions, edits, cancellations and executions — that arrive with heavy-tailed inter-event times and can trigger rapid regime shifts. Prior generative approaches modeled these messages autoregressively by splitting each message into many sub-tokens and simulating a clearing house step by step, which makes inference latency scale linearly with the prediction horizon and makes the models task-specific and slow.
The goal of this paper is to build a general-purpose, encoder-only foundation model for LOB messages that can be pre-trained once and then fine-tuned cheaply for many downstream tasks, while using far shorter context than previous methods.
Key Contributions
-
A one-token-per-message tokenizer. Each message (side, message type, coarse price and volume buckets, plus a volume "on-level" indicator) becomes a single discrete token, combined with continuous, piecewise-linear-geometric-scaled price, volume and time values. This yields 293 distinct message tokens from the training data and avoids the token explosion of prior models that decompose each message into dozens of sub-tokens.
-
Continuous-time rotary attention plus a leakage-safe pre-training objective. Continuous Rotary Position Embedding (RoPE) operates on cumulative time differences between messages rather than integer positions, and the Masked Message Modeling (MMM) objective jointly masks messages and their surrounding order book snapshots (snapshots masked for 90% of positions randomly) to prevent label leakage.
-
A hybrid discrete–continuous prediction head with a "Combined" inference scheme. A token classification head is paired with three regression heads that take dual input (token prediction logits concatenated with final hidden states); at inference the predicted token's quantization range bounds the continuous regressor. The paper reports this beats token-only and regressor-only inference on distributional fidelity.
-
Efficiency and transferability. Because LOBERT predicts a single token per message, the authors report the effective sequence length drops by roughly an order of magnitude on their data, and the pre-trained model fine-tunes efficiently to diverse downstream tasks with lightweight task-specific heads.
Main Findings
-
Next-message prediction beats the S5 baseline on every message component. Comparing LOBERT (1.1M parameters) against a reconstruction of the S5 model (1.2M parameters, no Book Module) trained on the same data with roughly equal training time: full-message accuracy 26.4% versus 6.1%; message type 60.6% versus 50.8%; side or direction 65.1% versus 52.2%; quantized price 51.8% versus 18.6%; quantized volume 72.1% versus 32.1%.
-
The order book snapshot module adds a further gain. LOBERT with Book Module achieves 61.9% message type, 65.4% side/direction, 53.8% quantized price, 72.2% quantized volume, and 27.8% full-message accuracy.
-
Predicted distributions track the real ones reasonably well. Pearson correlations between predicted and real distributions on the test data are 0.55 for price, 0.37 for volume, and 0.52 for time. The authors note a deviation at price level 10, hypothesized to come from conflict between the token head and the price regression head.
-
The Combined inference mode gives the best distributional fidelity. On the test split, measured by Wasserstein-1 (W1), Jensen–Shannon divergence (JSD) and Total Variation distance (TVD), Combined achieves Price W1 10.04 / JSD 0.2111 / TVD 0.1839 and Volume W1 76.86 / JSD 0.1696 / TVD 0.1200, versus Token only (Price 15.44 / 0.5566 / 0.6134; Volume 85.36 / 0.2446 / 0.1585) and Regressor only (Price 10.85 / 0.3959 / 0.4767; Volume 86.09 / 0.6623 / 0.8395). Time difference is predicted by regressors only (W1 10.658, JSD 0.2681, TVD 0.3270) and is included for reference.
-
LOBERT matches or beats DeepLOB on mid-price direction. Across horizons H = 10, 50 and 100 messages and confidence thresholds from 0.3 to 0.9, LOBERT's macro-F1 and coverage are consistently equal to or better than a re-trained DeepLOB. At confidence threshold 0.3, LOBERT scores F1 0.51 / coverage 1.00 (H=10), 0.56 / 1.00 (H=50) and 0.55 / 1.00 (H=100), against DeepLOB's 0.44 / 1.00, 0.53 / 1.00 and 0.55 / 1.00 respectively.
-
LOBERT responds more sharply to confidence filtering. As the confidence threshold rises from 0.3 to 0.9 at H = 100, macro-F1 rises from 0.55 to 0.88 while coverage falls from 1.00 to 0.10, indicating better ranking of easy versus hard cases than DeepLOB, whose confidence is described as more uniform.
-
Inference is currently a weakness. LOBERT runs at 281.87 predictions per second on a single NVIDIA V100 GPU with batch size 1, about 47% of DeepLOB's throughput, i.e. slower by 53%. The authors present quantization (up to 4x), structured pruning (up to 2x), early-exit strategies (roughly 1.57x), efficient attention (up to 1.5x) and continual-inference reformulations (up to two orders of magnitude) as untapped levers.
-
Data and training setup. Training used 80 trading days of Nasdaq ITCH feed data on AAPL, INTC, MSFT and FB between 2015.05.11 and 2015.09.01, 10 validation days between 2015.09.02 and 2015.09.16, and the final 10 days from 2015.09.17 to 2015.09.30 for testing. The dataset contains 470M messages split into 919k non-overlapping sequences of 512 messages each. The authors note that data access limits meant only millisecond-level time differences were used rather than nanoseconds, so the time aspect is only partially covered.
Methodology in Plain English
The researchers treat the stream of order book events as if it were a language, and adapt BERT to read and predict it.
Turning messages into tokens. Each message's price is first converted into a distance in ticks from the best opposing quote (the best bid for a sell order, the best ask for a buy order). That distance is bucketed into levels (0, 1, 2, 3, 5, 10), and a duplicate is passed through a scaling function that grows linearly for small values and then geometrically toward an asymptote, so rare extreme values are not lost. Volume is handled the same way with bucket levels (0, 50, 100, 200) — chosen because over 60% of all volume values in the dataset are exactly 100 units — plus a binary flag for whether the volume sits exactly on a bucket. Time is stored only as a continuous scaled inter-arrival difference, since no clear repeating pattern was found beyond heavy tails. Side, message type, bucketed price, bucketed volume and the volume flag are glued together with a colon delimiter to form one discrete token per message, producing 293 message tokens plus special padding, masking and unknown tokens.
Adding context. Every message is accompanied by a snapshot of the order book immediately after it: 40 values covering the 10 best bid prices with volumes and the 10 best ask prices with volumes, with volumes squashed into [0,1) by an exponential transform and prices expressed as tick distances from the opposing best quote, clipped and normalized.
The model. LOBERT merges four embedding sources additively — the discrete message token, continuous price difference, continuous volume, and optionally the snapshot through a learned gating mechanism — then adds learned positional embeddings and continuous Rotary Position Embedding computed on cumulative time, so attention is aware of real elapsed time, not just order. The encoder is a standard stack of multi-head self-attention and feed-forward blocks with Gelu activations, layer normalization, residual connections and dropout of 0.1. Outputs fan out into one classification head for the next token (cross-entropy loss) and three regression heads for price, volume and time (weighted MSE losses), where each regression head also sees the token logits.
Two-stage training. First, Masked Message Modeling: a bidirectional LOBERT reconstructs randomly masked messages, with snapshots also masked for 90% of positions so the answer cannot leak in through the snapshot. Regression losses are balanced after the first epoch while token loss is weighted higher. Second, fine-tuning: causal next-message prediction (with a triangular mask) and mid-price direction classification. Optimization uses AdamW with a learning rate of 5×10⁻⁵, weight decay 0.01 (disabled for bias and layer norm parameters), Cosine Annealing with Warm Restarts (T₀ = 40,000 steps, T_mult = 2, minimum learning rate 5×10⁻⁶), 10 epochs, batch size 32, and validation every 15,000 steps.
Why This Matters
Impact on research. The paper argues that a single pre-trained encoder can serve as a reusable foundation model for LOB data, in the way BERT changed language modeling: instead of building a bespoke generative simulator per task, practitioners could pre-train once and attach lightweight heads. It also reframes evaluation — the paper introduces a distributional-fidelity comparison (W1, JSD, TVD) of predicted versus true marginals for price, volume and time, and demonstrates selective prediction with confidence thresholds, which is a natural fit for trading systems that can abstain.
Real-world applications (as identified in the paper):
- Market making, where quotes must react to message flow on millisecond time scales.
- Arbitrage strategies that depend on rapidly detecting short-horizon mispricing.
- Market impact modeling and optimal execution, where a realistic generator of order flow supports execution scheduling.
- Spoofing detection and other surveillance applications (named in the project the authors thank the Nasdaq Nordic Foundation for funding).
Industry relevance. The paper is explicit that high-frequency trading requires forecasts within milliseconds, and that prior autoregressive message generators were too slow because latency scaled with the horizon. The claim that LOBERT reduces effective sequence length by roughly an order of magnitude and supports fine-tuning for arbitrary horizons speaks directly to that constraint — though the authors also acknowledge LOBERT's own current inference speed is 47% of DeepLOB's throughput on an NVIDIA V100, making latency an open engineering problem rather than a solved one.
Future Directions
- Validating long-horizon generative quality. The authors state the current experiments do not guarantee the model can produce realistic long-term sequences and plan extensive sequence-level analysis.
- Broadening the fine-tuning task coverage. The breadth of task types LOBERT can be adapted to is explicitly described as an unfinished study.
- Tracking order state. Individual order IDs are not modeled, and the authors note LLMs are known to struggle with tracking entity state through changes — LOBERT may face the same difficulty in tracking the size and location of orders when price levels move.
- Reducing inference cost. The paper lists quantization, structured pruning, early-exit strategies, efficient attention, and continual-inference reformulations (Continual Transformers / Deep Continual Transformers) as routes to higher throughput and, potentially, redundancy-free continual operation.
- Enriching inputs. The roadmap includes feeding in market state information such as time of day, the current bid-ask spread, and full LOB snapshots.
Target Audience
Quantitative researchers and machine learning engineers working on market microstructure and high-frequency trading; NLP and transformer researchers interested in domain adaptation to non-language event streams; and practitioners in execution, market making or surveillance who need a pre-trainable, fine-tunable model of order flow. Readers need a working knowledge of transformer architecture and limit order book mechanics to get the most out of it; the paper is less suited to readers seeking a gentle introduction to either.
Authors’ abstract
Modeling the dynamics of financial Limit Order Books (LOB) at the message level is challenging due to irregular event timing, rapid regime shifts, and the reactions of high-frequency traders to visible order flow. Previous LOB models require cumbersome data representations and lack adaptability outside their original tasks, leading us to introduce LOBERT, a general-purpose encoder-only foundation model for LOB data suitable for downstream fine-tuning. LOBERT adapts the original BERT architecture for LOB data by using a novel tokenization scheme that treats complete multi-dimensional messages as single tokens while retaining continuous representations of price, volume, and time. With these methods, LOBERT achieves leading performance in tasks such as predicting mid-price movements and next messages, while reducing the required context length compared to previous methods.