Research
Online Domain-aware LLM Decoding for Continual Domain Evolution
Online Domain-aware LLM Decoding for Continual Domain Evolution Overview Research area: Machine learning / natural language generation, specifically inference-time adaptation and decoding strategies f

- arXiv
- 2602.08088
- Published
- 2026-02-08
- Authors
- Mohammad Abu-Shaira, Weishi Shi
AI summary
Online Domain-aware LLM Decoding for Continual Domain EvolutionOverview
- Research area: Machine learning / natural language generation, specifically inference-time adaptation and decoding strategies for large language models under concept drift.
- Technical level: Intermediate. The paper assumes familiarity with next-token distributions, softmax/logits, temperature scaling, ROUGE/BLEU/BERTScore, and divergence measures such as Jensen–Shannon Divergence.
- Scope: The paper proposes ODD (Online Domain-aware Decoding), an inference-time decoding framework that fuses a base LLM's next-token distribution with an online prefix-tree (Trie) prior weighted by confidence, disagreement, and temporal continuity signals, and evaluates it on syntactically and semantically drifted versions of a single telco chatbot dataset using gpt2-medium.
What This Paper Is About
LLMs are usually adapted to a domain by offline fine-tuning, which assumes the domain is static. In practice, domain knowledge keeps changing through new products, regulations, services, and interaction patterns, and retraining for every change is computationally infeasible; this produces concept drift that degrades prediction quality. The goal of this work is to adapt the LLM's next-token distribution in real time without updating model weights, without auxiliary models, and without external retrieval.
Key Contributions
- ODD, an inference-time online decoding framework. ODD performs probability-level fusion between a base LLM and a prefix-tree prior, guided by adaptive confidence modulation using disagreement and continuity signals, with claims of no retraining, no auxiliary models, and no external retrieval.
- An online prefix tree (Trie) as a drift-aware domain prior. The Trie is continuously updated with an n-gram scheme inserting token sequences up to length N, and each node stores three features — Frequency (F), Length (L), and Recency (R) — that are combined with configurable weights into candidate scores converted to a sparse probability distribution.
- A calibration and confidence-modulation mechanism. Adaptive temperature scaling matches the LLM's peak probability to the Trie's peak probability before fusion, and the interpolation weight is derived from normalized-entropy LLM confidence and maximum-probability Trie confidence, adjusted by top-k disagreement (k=5) and a temporal continuity term based on consecutive agreement steps.
- An empirical evaluation across three drift scenarios. Using the Bitext Telco LLM Chatbot Training Dataset (approximately 26,000 samples, 26 intents) with placeholders substituted to construct abrupt, incremental, and gradual drift, ODD is benchmarked against LLM-Greedy and LLM-Temp Scaled.
Main Findings
- ODD outperforms both baselines on all reported NLG metrics. Under abrupt, incremental, and gradual drift, ODD exceeds LLM-Greedy and LLM-Temp Scaled across Exact Match, Edit Distance, BLEU, ROUGE-L, Cosine Similarity, ChrF, and BERTScore in Table 1.
- Semantic alignment gains are the largest. Under abrupt drift, ODD reaches Cosine Similarity 0.968 versus 0.865 (Greedy) and 0.863 (Temp Scaled), and BERTScore 0.928 versus 0.911 and 0.907.
- Incremental drift yields the strongest stability. ODD achieves Cosine Similarity 0.976 versus 0.865 and 0.862, BERTScore 0.935 versus 0.908 and 0.902, and the best ChrF of 83.83 versus 78.30 and 76.89.
- Gradual drift produces the headline gains. ODD attains ROUGE-L 0.825, an absolute gain of 0.065 over the strongest baseline, and Cosine Similarity 0.971, a relative improvement of 13.6% over the best baseline; ChrF is 82.82 versus 74.03 and 72.71.
- Exact Match is non-zero only for ODD. ODD scores 0.096 (abrupt), 0.052 (incremental), and 0.037 (gradual) Exact Match, while both baselines score 0 across all three scenarios.
- Drift severity affects all methods. Baseline scores decline from abrupt to incremental to gradual drift (for example, Greedy ChrF drops from 80.374 to 78.304 to 74.031), and ODD declines as well while remaining highest (83.147 to 83.826 to 82.824).
- Qualitative adaptation to new entities. After abrupt drift, baselines keep generating outdated Concept 1 templates such as "To sign up for a Mobile Voice and Data Communication Service…", whereas ODD produces "To sign up for a Talk + Net Packet plan with TalkNow Crew…".
- Ablation results. Using disagreement (Ω_t) and continuity (Γ_t) jointly produced the most stable behavior; enabling only one signal reduced robustness. Disabling temperature calibration caused confidence-scale mismatches between the LLM and Trie distributions. Trie feature weights were fixed to uniform values (λ_F = λ_L = λ_R = 1/3).
- Low overhead. Calibration and confidence computation add less than 0.2 ms overhead per decoding step, and Trie updates and retrievals are reported as sub-millisecond per decoding step.
Methodology in Plain English
The method keeps an online dictionary of token sequences seen so far, organized as a prefix tree. When the model generates text, it looks up every suffix of the current prefix in that tree and collects the tokens that could come next. Each candidate token gets a score built from three quantities: how often it appeared (Frequency, compressed with log(1+F) and normalized), how deep the matching phrase is relative to the current prefix length (Length, normalized by |π_t|), and how recently it appeared (Recency, modeled as exp(−Δ/Δ_max)). A weighted sum of these three normalized values, with non-negative weights summing to 1, gives the candidate score, which is turned into a sparse probability distribution using a top-preserving normalization that keeps the maximum score as the peak.
Because the neural distribution from the LLM and this statistical distribution sit on different confidence scales, the LLM's logits are re-scaled by an adaptive temperature chosen so that the LLM's maximum probability exactly equals the Trie's maximum probability; the paper notes a unique solution always exists because the peak softmax value is continuous and strictly monotonic in T, and it is found by a 1D monotonic root-finding method (bisection). The two distributions are then blended with a convex combination.
The blending weight is not fixed. The LLM's confidence is measured as one minus its normalized entropy (c_LM = 1 − H/H_max, with H_max = log|V|), and the Trie's confidence is the maximum probability of its top-preserving candidate. Two context signals then adjust these: disagreement Ω_t, computed from the Jensen–Shannon Divergence between the two distributions' top-5 tokens (Ω_t = min(1, sqrt(JSD(p,q))), which penalizes the LLM via c'_LM = c_LM(1 − Ω²)), and continuity Γ_t = 1 − exp(−r_t/3), where r_t counts how many consecutive steps both experts picked the same top token, which amplifies the Trie via c'_trie = c_trie + (1 − c_trie)c_trie²Γ. The final weight is γ_t = c'_LM/(c'_LM + c'_trie), and the next token is the argmax of the mixed distribution. Insertion into the Trie is O(L), all-suffix retrieval is O(L²), worst-case memory is O(L_total), and empirical memory is described as O(U) for U distinct placeholder prefixes; both insertion and retrieval are reported as independent of the number of stored sequences N.
Why This Matters
- Research impact: The paper pushes adaptation into the decoding step rather than the training step, arguing that logit-control methods need auxiliary controllers or extra passes, prefix/context methods use rigid or static priors, retrieval-guided decoding adds latency and depends on index freshness, and parameter-efficient adaptation plus model editing require offline training or permanent weight changes. It also claims to be the first inference-time, online domain-aware decoding framework for mitigating concept drift without external retrieval or additional training.
- Customer support assistants: A telco chatbot must handle newly launched plans and brands without being retrained, which is exactly the scenario tested with the Bitext dataset.
- Regulatory and compliance-driven domains: When rules or policies change, the Trie can absorb new terminology or placeholders online while the base model's weights stay untouched.
- Product and pricing catalogs: Frequently updated product names, account identifiers, and service tiers can be represented as tokens the Trie tracks for recency and frequency.
- Latency-constrained deployments: Because no external retrieval index and no per-token auxiliary forward/backward passes are required, and reported overhead is under 0.2 ms per decoding step, the approach targets real-time serving.
- Industry relevance: The framework offers a zero-training update path for deployed LLMs, avoiding the cost, catastrophic-forgetting risk, and data-privacy/regulatory concerns the paper associates with repeated fine-tuning or retraining.
Future Directions
- Alternative Trie feature weightings. The paper fixes λ_F = λ_L = λ_R = 1/3 and explicitly states that exploring alternative weightings remains future work, leaving the ability to prioritize recency, length, or frequency untested.
- Generalization beyond placeholder drift and one dataset. Evaluation is limited to the Bitext Telco dataset, where placeholders are substituted to simulate drift; testing on other domains and on non-placeholder drift would clarify how broadly the approach transfers.
- Scaling to larger base models. Experiments use only gpt2-medium (355M parameters); whether the calibration and confidence signals behave the same for much larger LLMs is not reported.
- Sensitivity to key hyperparameters. The top-k value (fixed at k=5 in all experiments), the n-gram insertion depth N, and the continuity constant (the "3" in Γ_t) are not ablated in the provided content, leaving their impact as an open question.
Target Audience
Researchers and practitioners working on LLM inference, decoding strategies, and continual or online adaptation will get the most from this paper, particularly those interested in concept drift, non-parametric memory priors, and training-free deployment. It is also relevant to applied engineers building customer-support or domain-specific assistants who need updates without retraining, and to readers who want a compact example of combining a neural distribution with a statistical prior through confidence-based interpolation. Some background in probability distributions, entropy, and NLG evaluation metrics is assumed.
Authors’ abstract
LLMs are typically fine-tuned offline on domain-specific data, assuming a static domain. In practice, domain knowledge evolves continuously through new regulations, products, services, and interaction patterns. Retraining or fine-tuning LLMs for every new instance is computationally infeasible. Additionally, real-world environments also exhibit temporal dynamics with shifting data distributions. Disregarding this phenomenon, commonly referred to as concept drift, can significantly diminish a model's predictive accuracy. This mismatch between evolving domains and static adaptation pipelines highlights the need for efficient, real-time adaptation without costly retraining. In response, we introduce Online Domain-aware Decoding framework (ODD). ODD performs probability-level fusion between a base LLM and a prefix-tree prior, guided by adaptive confidence modulation using disagreement and continuity signals. Empirical evaluation under diverse drift scenarios demonstrates that ODD consistently surpasses LLM-Greedy and LLM-Temp Scaled across all syntactic and semantic NLG metrics. It yields an absolute ROUGE-L gain of 0.065 and a 13.6% relative improvement in Cosine Similarity over the best baseline. These results demonstrate ODD 's robustness to evolving lexical and contextual patterns, making it suitable for dynamic LLM applications.