Research
OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning
OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning Overview Research area: Machine learning — time-series foundation models, time-series langua

- arXiv
- 2609.40265
- Published
- 2026-09-30
- Authors
- Tony Chen, Timo Stoffregen, Maxwell Xu, Thomas Kaar, Martin Maritsch, Geremia Pompei, Nicolas Zumarraga, Robert Jakob, Paul Schmiedmayer, Patrick Langer, Juncheng Liu
AI summary
OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and ReasoningOverview
Research area: Machine learning — time-series foundation models, time-series language models (TSLMs), and parameter-efficient expert composition (LoRA mixture-of-experts).
Technical level: Advanced. The paper assumes familiarity with time-series forecasting metrics (MASE, CRPS/RCRPS), quantile forecasts, LoRA adapters, and mixture-of-experts routing.
Scope: The paper introduces OpenTSLM TeeMoE, a generalist time-series language model built from three independently trained low-rank experts composed by a learned controller over a shared Qwen3.6-27B backbone, and evaluates it on GIFT-Eval, Context is Key, and TimeSeriesExam.
Publication details: arXiv:2609.40265v1 [cs.LG], 30 Sep 2026, CC BY 4.0. Authors are affiliated with Columbia University, Stanford University, Aionic Labs, Google, Agentic Systems Lab at ETH Zürich, University of Pisa, and the National University of Singapore; the paper notes shared last authors.
What This Paper Is About
Real time-series applications need models that can do three different things: produce numerical forecasts, incorporate textual context that changes what the future should look like, and answer analytical questions about observed signals. Today these capabilities live in separate model families — numerical specialists forecast best, while language-based models understand context and reasoning better — and combining them usually means losing the accuracy that made the specialists useful.
The paper's central question is whether a general-purpose model can acquire all three capabilities without diluting specialist performance. The authors answer by training each capability separately as its own low-rank adapter, freezing them, and training a small controller that decides, per request, how to weight and route those adapters.
Key Contributions
-
A generalist TSLM built from independently trained experts over a shared backbone. OpenTSLM TeeMoE combines a foundation-model aggregation expert, a native forecasting expert, and a temporal analysis expert within one Qwen3.6-27B backbone. The authors state this is, to their knowledge, the first use of MoE to construct a generalist TSLM, and they release code and model checkpoints publicly.
-
Competitive performance across three capability benchmarks. A single composed checkpoint ranks among the top three on GIFT-Eval (mean MASE rank 19.990), Context is Key (RCRPS 0.115), and TimeSeriesExam (78.552% accuracy), as of September 25, 2026 — benchmarks where specialized models usually dominate.
-
A study of how MoE composition affects TSLM capabilities and training. The paper compares individual experts, alternative composition strategies (top-1 routing, equal weights, full-strength adapters), and a jointly trained baseline, characterizing capability preservation, expert interaction, and training cost.
-
A language-refined numerical aggregation design. The aggregation expert exposes both an ensemble prediction and its constituent candidate distributions to the language backbone, extending prior forecast-combination work by refining the reference forecast rather than only selecting among candidates.
Main Findings
-
Top-three placement across three benchmarks. The same composed checkpoint achieves a mean MASE rank of 19.990 on GIFT-Eval, an RCRPS of 0.115 on Context is Key, and 78.552% accuracy on TimeSeriesExam. GIFT-Eval evaluation covers 97 evaluation cells and 371,330 forecast windows; CiK covers 355 instances across 71 task types; TSE uses 746 questions in its official v1.1 release.
-
Broad task coverage without losing predictive strength. TeeMoE outperforms the selected individual numerical foundation models on GIFT, achieves lower CiK RCRPS than the locally evaluated time-series language models under the documented transfer protocols, and leads the TSE comparison.
-
Smaller total size than several larger comparators. TeeMoE uses approximately 40B total parameters, including a 27B language backbone, external forecasting models, and learned adapters, encoder, and decoder. Despite this, it outperforms Llama-3.1-405B-Instruct (405B parameters) on CiK, the GPT-oss-120B hybrid agent (approximately 117B parameters) on TSE, and GPT-5.2-based forecast correction on CiK.
-
Improvement over the unadapted backbone. Unadapted Qwen3.6-27B is already a strong starting point, but TeeMoE reduces CiK RCRPS by 0.036 and raises TSE accuracy by 3.619 percentage points with the same inputs.
-
Complementary specialist strengths. The native forecasting expert gives the best specialist CiK score (0.123), while the analysis expert leads on TSE (78.418%). Composition reduces RCRPS by 6.6% relative to the native expert, improves TSE accuracy by 0.134 percentage points over the analysis expert, and slightly improves the aggregation expert's GIFT mean rank.
-
Expert interactions add only small predictive gains. Top-1 adapter routing keeps comparable performance, with a noticeably smaller CiK improvement than soft composition and slightly worse GIFT and TimeSeriesExam results — suggesting capability preservation extends across both soft composition and single-expert selection.
-
Learned weighting beats fixed mixtures. Equal weights (1/3) slightly improve GIFT rank, worsen CiK, and reduce TSE accuracy by 0.402 percentage points; full-strength adapters improve GIFT rank but worsen CiK and reduce TSE accuracy by 2.949 percentage points.
-
Separate specialization beats joint training on task metrics. TeeMoE achieves a larger CiK reduction (0.036 versus 0.028), a larger GIFT rank reduction (0.670 versus 0.608), and a 0.670-percentage-point TSE gain over the joint baseline, which uses one shared adapter and a separate output selector.
-
Separate specialization is cheaper. Separate specialization costs 30.756 H100 GPU-hours versus 58.657 for the authors' joint recipe, excluding shared numerical preparation.
-
Learned weighting matters for numerical ensembling. On GIFT, learned XGBoost weighting improves substantially over uniform pooling, while adding more models with equal weights does not help. Refinement further improves the already strong blend, though the authors note rank gains do not directly establish which component contributes more.
-
Context is the main CiK driver. History-only ensembles score 0.283–0.292 RCRPS on CiK, well behind TeeMoE's 0.115.
Methodology in Plain English
-
Pick a shared language backbone. All capabilities sit on one frozen pretrained model, Qwen3.6-27B.
-
Train three experts separately, each with its own data and objective.
- The aggregation expert first builds a reference forecast by ensembling pretrained numerical forecasters — an XGBoost regressor predicts relative ranks of eight selected models and converts them into weights for pooling their cumulative distribution functions, and this is blended with a scalar weight against Toto-FnF, a public ten-model ensemble. The two pools share five forecasters, giving a union of thirteen candidates. The forecast was fitted on 576,920 forecasting windows across economics, energy, healthcare, environmental monitoring, retail, transport, and computing systems, and the aggregation adapter, numerical encoder, and forecast decoder were jointly trained on 4,096 examples from the GIFT training split using a weighted SmoothL1 loss plus a squared-edit penalty.
- The native forecasting expert generates future timestamp–value pairs directly from history and textual context, with no external forecasters. It trains for one epoch on 20,000 contextual forecasting examples with teacher-forced cross-entropy and forward KL regularization to the frozen base model.
- The temporal analysis expert answers analytical questions — patterns, anomalies, series comparisons, statistical relationships — training for one epoch on 12,000 time-series question-answering examples with forward and reverse KL regularization.
-
Convert forecasts into language-model-readable tokens. A numerical encoder summarizes reference quantiles and the thirteen candidate distributions at eight evenly spaced horizon positions using a 53-value feature vector per position (nine quantile levels from 0.1 to 0.9, candidate medians, widths and asymmetries, position and scale coordinates). Projections combine these into a 5,120-dimensional embedding per position, and the eight embeddings appear three times in a 24-token sequence.
-
Refine rather than replace. A zero-initialized decoder predicts a bounded, dimensionless correction, adjusted by a disagreement gate that shrinks the edit when candidate forecasts disagree more than the reference uncertainty warrants. The same shift is applied to every quantile at a time step, preserving quantile order and width.
-
Freeze everything, train a controller. The backbone, all three adapters, and aggregation components stay frozen. A linear controller maps a request representation to three softmax weights shared across all adapted layers and fixed for the whole response. Composition is a weighted sum of the experts' low-rank updates added to the frozen weight matrices. The controller trains for one epoch on 1,000 examples, combining capability prediction losses with a numerical-versus-textual output-format loss.
-
Route the output per request. If the aggregation weight is above 0.5, the request goes through the numerical forecast decoder; otherwise it is generated as language tokens. No context/no-context switch is hard-coded. Execution uses a two-pass scheme drawn from X-LoRA: an adapter-disabled pass forms the request representation, and a second pass applies the mixed adapters.
Why This Matters
-
Research impact. The paper offers evidence that modular expert composition can produce a strong generalist time-series model that preserves the strengths of independently trained specialists — and reports that this route was also cheaper than the authors' joint-training recipe (30.756 versus 58.657 H100 GPU-hours). It also provides a reusable comparison framework across three capability benchmarks rather than one.
-
Real-world applications (as described or implied in the paper):
- Demand planning: forecast future demand, predict the effect of a planned intervention, and interpret an unusual historical pattern.
- Healthcare: forecasting and temporal interpretation in a domain the ethics statement specifically names as consequential.
- Finance: temporal interpretation and forecasting where benchmark accuracy is explicitly stated not to establish fitness for autonomous decisions.
- Public infrastructure: the third consequential domain named in the ethics statement.
-
Industry relevance. The design lets practitioners develop or revise one capability without retraining the others — an expert and the controller can be retrained while the rest stay fixed. The model also interfaces with existing forecasting infrastructure: it consumes outputs from thirteen candidate models rather than replacing them.
-
Caveats the authors state. Benchmark accuracy does not establish fitness for autonomous decisions in healthcare, finance, or public infrastructure; users should validate on the intended domain and retain human oversight.
Future Directions
-
Workflows that alternate between forecasting and reasoning. The authors describe extending evaluation to tasks that interpret a forecast and revise it in response to additional context as a natural next step.
-
Testing beyond one backbone family. The paper states that extending the study beyond a single backbone family would clarify how these findings carry over to other models and training recipes.
-
Better expert interaction. Because learned routing contributes only small additional predictive gains over top-1 selection, the value of simultaneous adapter contributions remains an open question the authors flag for future generalist TSLMs.
-
Broader benchmark coverage. The study prioritizes GIFT-Eval, Context is Key, and TimeSeriesExam because they provide established protocols and published specialist comparisons; the authors note these benchmarks assess capabilities separately rather than in combination.
Target Audience
- Machine learning researchers working on time-series foundation models, time-series language models, and mixture-of-experts or LoRA composition methods.
- Practitioners building production forecasting systems who need one model to serve forecasting, context-conditioned prediction, and analytical question answering.
- Engineers evaluating parameter-efficient adaptation: the paper's cost accounting (30.756 versus 58.657 H100 GPU-hours) and revision workflow are directly relevant to teams deciding between modular and joint training.
- Benchmark and evaluation researchers interested in how generalist models are compared against task-specific specialists across GIFT-Eval, Context is Key, and TimeSeriesExam.
- Readers tracking reproducibility infrastructure will find the released code and Hugging Face checkpoints useful, along with the documented training recipes, evaluation protocols, and the paper's AI Use Statement describing how generative AI tools were and were not used.
Authors’ abstract
Real-world time-series applications increasingly require models that can handle time series forecasting, context-conditioned prediction, and language-based temporal reasoning. Yet current time-series foundation models remain fragmented across these capabilities: numerical specialists often provide the strongest forecasts, while language-based models offer broader contextual understanding and analysis. A central challenge is to unify these heterogeneous capabilities without reducing their individual performance. We introduce OpenTSLM TeeMoE, a generalist time-series language model that can forecast directly from observed time series, reason over textual context and temporal patterns, and synthesize and refine predictions from external numerical forecasting specialists. We independently train three low-rank experts for forecast aggregation, native forecasting, and temporal analysis over a shared backbone. A learned LoRA mixture-of-experts controller then weights their frozen parameter updates for each request. Our proposed model achieves strong performance on widely used benchmarks for time series forecasting, context-conditioned prediction, and language-based temporal reasoning, ranking among the top three on GIFT-Eval by mean MASE rank, Context is Key by RCRPS, and TimeSeriesExam by accuracy.