Research
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
Overview Research area: Efficient large language model inference — specifically speculative decoding, a family of techniques for accelerating autoregressive generation. Technical level: Intermediate.

- arXiv
- 2609.15504
- Published
- 2026-09-14
- Authors
- Ilya Koziev, Leonid Sinev, Ivan Oseledets
AI summary
Overview
Research area: Efficient large language model inference — specifically speculative decoding, a family of techniques for accelerating autoregressive generation.
Technical level: Intermediate. The paper assumes familiarity with autoregressive decoding, KV caches, floating-point precision formats (BF16/FP32), and standard evaluation harnesses, but its central argument is conceptually simple.
Scope: An independent reproduction of the Orthrus hybrid autoregressive–diffusion decoding architecture that tests its "lossless" claim by comparing generated token trajectories against the reference autoregressive model under BF16 and FP32 precision.
What This Paper Is About
Orthrus (Nguyen et al., 2026) augments a frozen autoregressive language model with a lightweight diffusion view that predicts several future tokens in parallel, then uses an intra-model consensus mechanism to validate those tokens. The authors claim this process is strictly lossless: it should produce exactly the same output sequence as the original autoregressive model. This paper independently reproduces Orthrus and asks whether that exact equivalence actually holds when the model is run in finite-precision arithmetic rather than in an idealized mathematical setting.
Key Contributions
- An independent implementation of Orthrus training and inference, evaluated side by side with the authors' released checkpoint (chiennv/Orthrus-Qwen3-1.7B).
- A demonstration that a model trained on teacher-generated, on-policy distillation data (Orthrus-1.7B-final) achieves competitive or higher Tokens Per Forward (TPF) than the released checkpoint in most evaluated domains — higher in 10 of 12 domains, with the checkpoint better in the remaining two.
- Evidence that exact sequence-level equivalence is highly sensitive to numerical precision: only 45% (authors' checkpoint) and 43% (independently trained model) of trajectories match under BF16, while FP32 yields exact matching on all 1,190 evaluated prompts.
- An analysis showing that trajectory divergence is systematically associated with higher response-conditional perplexity under the reference model, and that divergence does not translate into systematic downstream benchmark degradation.
Main Findings
-
BF16 trajectory matching is far from perfect. Across 1,190 prompts spanning 12 domains, the authors' Orthrus-Qwen3-1.7B reproduced the Qwen3-1.7B trajectory exactly in 0.45 ± 0.03 of cases, and the independently trained Orthrus-1.7B-final in 0.43 ± 0.03 of cases. Diverging-trajectory rates were 0.55 ± 0.03 and 0.57 ± 0.03 respectively.
-
Match rates vary strongly by domain. The authors' checkpoint reached 0.88 ± 0.07 on gec-en and 0.73 ± 0.09 on qa-ru, but only 0.11 ± 0.06 on poetry-en and 0.12 ± 0.06 on creative-en. The independently trained model's lowest domain was poetry-en at 0.08 ± 0.05.
-
Divergence tracks reference-model perplexity. Mean response-conditional perplexity was 1.11 ± 0.01 for matching trajectories and 1.28 ± 0.01 for diverging ones (authors' checkpoint); 1.10 ± 0.01 and 1.28 ± 0.01 for the independent model. A logistic regression controlling for response length and domain gave β₁ = −8.10 (95% CI [−10.50, −5.71], p = 3 × 10⁻¹¹) for chiennv/Orthrus-Qwen3-1.7B and β₁ = −10.92 (95% CI [−13.60, −8.24], p = 1.3 × 10⁻¹⁵) for Orthrus-1.7B-final.
-
Downstream benchmark scores do not degrade. On lm-eval-harness, Orthrus-1.7B-final had higher point estimates than the Qwen3-1.7B baseline on all three benchmarks: GSM8K exact match (flexible) 0.4147 ± 0.0136 versus 0.4003 ± 0.0135; HumanEval pass@1 0.4146 ± 0.0386 versus 0.4024 ± 0.0384; IFEval prompt-level loose accuracy 0.2181 ± 0.0178 versus 0.2015 ± 0.0173, and strict accuracy 0.1830 ± 0.0166 versus 0.1682 ± 0.0161. The authors state these differences should not be read as statistically significant improvements given the reported uncertainty.
-
FP32 restores exact equivalence. Repeating the same evaluation in FP32 gave a sequence match rate of 1.00 ± 0.00 and a diverging-trajectory rate of 0.00 ± 0.00 for both Orthrus variants across all 1,190 prompts.
-
TPF is competitive despite the independent training setup. Orthrus-1.7B-final achieved higher TPF than the released checkpoint in 10 of 12 domains, for example 3.44 ± 0.13 versus 3.24 ± 0.20 on code and 3.93 ± 0.13 versus 3.59 ± 0.22 on code-ru. The released checkpoint was better on gec-en (5.00 ± 0.30 versus 4.23 ± 0.17) and math (8.08 ± 0.65 versus 4.99 ± 0.16).
-
Benchmark equality cannot establish inference equivalence. Two systems can obtain identical or statistically indistinguishable task scores while producing different token sequences; a small numerical perturbation can even move a generation toward a benchmark-preferred answer.
Methodology in Plain English
The researchers first re-trained Orthrus from scratch so they would not be limited to the released checkpoint. They built a distillation corpus by taking prompts from public HuggingFace datasets and recording the greedy-decoded responses of the frozen Qwen/Qwen3-1.7B model, keeping only prompts between 50 and 1,000 characters. The resulting dataset contains 4,113,358 prompt–response samples. Training ran for 1 epoch on eight NVIDIA H100 GPUs with CUDA 13.3.73, PyTorch 2.13.0 and Transformers 5.8.0, using an initial learning rate of 2 × 10⁻⁴, batch size 10, cross-entropy loss, block size 8, 32 blocks, and a maximum sequence length of 3,072 tokens.
For evaluation, they assembled 12 text domains (code, creative writing, grammar correction, math, poetry, question answering, and translation, in English and Russian), with 100 prompts per domain except gec-en, which has 90 — giving 1,190 prompts in total. Generation used greedy decoding with max_new_tokens=128, do_sample=False, and temperature=0.0 under BF16 with eager attention on a single NVIDIA GeForce RTX 3090 (23 GB, compute capability 8.6), with Python 3.10.12, PyTorch 2.8.0+cu128, CUDA 12.8, and Transformers 5.8.1. They then compared token-by-token trajectories produced by each Orthrus variant against the plain Qwen3-1.7B reference, and repeated the whole comparison in FP32.
They also measured the response-conditional perplexity that Qwen3-1.7B assigns to each generated response, and fitted a logistic regression with exact match as the binary outcome and log perplexity, response length, and domain as predictors. Finally, they ran GSM8K, HumanEval, and IFEval through lm-eval-harness to see whether trajectory divergence shows up as task-level performance loss.
Why This Matters
Impact on research. The paper reframes what "lossless" means for neural decoding systems. An algorithm can preserve the intended autoregressive computation in principle while producing different discrete outputs when implemented with finite-precision arithmetic and a stateful KV cache. The authors argue that any losslessness claim should specify both the operational criterion for equivalence and the numerical precision under which it is measured. They also note that benchmark-level equality is not evidence of inference equivalence.
Real-world applications:
- Serving LLM inference with speculative or parallel decoding, where accuracy guarantees matter for reproducibility and auditing.
- Regulated or high-stakes generation pipelines where exact output reproducibility may be a compliance requirement.
- Benchmarking and evaluation practice, since higher task scores can arise from numerical noise rather than genuine model improvement.
- Reproducibility work for other acceleration methods that make "exact output" claims, such as speculative decoding with draft models.
Industry relevance. Orthrus is an inference-acceleration technique, and the paper is explicit that its findings do not undermine Orthrus's utility: the method still retains high fidelity while providing substantial acceleration, with TPF values ranging roughly from 1.58 to 8.08 across domains and models. The practical lesson for deployment is that precision settings are a design decision with observable behavioral consequences, not an implementation detail.
Future Directions
- Identify the specific computational stages — which layers or operations — are responsible for the precision-dependent deviations, which the authors explicitly leave to future work.
- Develop a more precise operational definition of "lossless" for neural decoding that accounts for finite-precision arithmetic and stateful KV caches.
- Investigate whether trajectory divergence under low precision can be mitigated without moving entirely to FP32, since FP32 matching comes at a computational cost the paper does not evaluate.
- Extend the precision analysis to other speculative and parallel decoding methods to see whether the effect generalizes beyond Orthrus.
- Relate the observed perplexity–divergence association to the broader literature on LLM inference reproducibility across precision and hardware configurations (Yuan et al., 2025).
Target Audience
Researchers and engineers working on LLM inference acceleration, speculative decoding, and parallel decoding architectures; practitioners responsible for deploying accelerated inference in settings where output reproducibility matters; and evaluation specialists who need to understand why matching benchmark scores does not establish that two systems compute the same thing. Readers need working familiarity with autoregressive decoding and floating-point formats, but the paper's central claim — that "lossless" depends on the arithmetic you run it in — is accessible to anyone who has debugged nondeterministic model output.
Authors’ abstract
Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by generating multiple tokens in parallel while using a frozen autoregressive backbone. Its central claim is that an intra-model consensus mechanism enables lossless speculative decoding, producing the same output sequence as the autoregressive model. We independently reproduce Orthrus and examine this claim under different numerical precisions. Under BF16 inference, exact trajectory matching occurs in only 45% of cases for the authors' checkpoint and 43% for our independently trained model across 1,190 prompts from 12 domains. The probability of exact matching is also strongly associated with the response-conditional perplexity of the reference model. Despite this trajectory divergence, Orthrus does not show systematic degradation on downstream lm-eval-harness benchmarks. In contrast, repeating the trajectory evaluation with FP32 yields exact trajectory matching on all evaluated prompts. These results show that the practical losslessness of Orthrus depends on numerical precision and that exact trajectory equivalence should be evaluated separately from downstream task performance.