Research
XLM: A Python package for non-autoregressive language models
Overview Research area: Natural Language Processing — software tooling for non-autoregressive language modeling. Technical level: Intermediate. The paper is a systems/tooling paper rather than a model

- arXiv
- 2512.17065
- Published
- 2025-12-18
- Authors
- Dhruvesh Patel, Durga Prasad Maram, Sai Sreenivas Chintha, Benjamin Rozonoyer, Andrew McCallum
AI summary
Overview
Research area: Natural Language Processing — software tooling for non-autoregressive language modeling.
Technical level: Intermediate. The paper is a systems/tooling paper rather than a modeling paper. It assumes familiarity with Python deep learning stacks (PyTorch, PyTorch Lightning, Hydra) and with basic language-modeling concepts such as training loops, losses, collators, and inference procedures.
Scope: The paper introduces XLM, a Python package (with a companion xlm-models package) for rapidly implementing, training, and systematically comparing small non-autoregressive language models.
What This Paper Is About
Autoregressive language models have mature, standardized libraries for training and inference, while non-autoregressive methods (masked diffusion, Gaussian diffusion, insertion, edit-based) are mostly implemented as one-off, bespoke codebases. Because each of these methods requires its own data collation, loss, and prediction logic, it is hard to reuse components or compare methods systematically. XLM is presented as a unified framework that makes implementing small non-autoregressive language models faster without sacrificing flexibility, and that can ship a suite of small pre-trained reference models.
Key Contributions
- A modular Python package (
xlm-core) built on PyTorch, PyTorch Lightning, and Hydra, providing shared, model-independent core components (HarnessandDataModule) while letting each model live in its own self-contained folder. - A design philosophy centered on "maximal independence," achieved through composition over inheritance, deliberate copy-over-branching of research code, and arbitrary runtime code injection via Hydra.
- A scaffolding workflow (
xlm-scaffold) that auto-generates the directory and file skeleton for a new model, plus hierarchical Hydra configuration composition for swapping model, loss, predictor, collator, and dataset components without editing Python. - A benchmark suite and companion
xlm-modelspackage that reproduces published results for ARLM, MDLM, MLM, and ILM, and is intended to grow into a set of reference implementations and small pre-trained models for the community.
Main Findings
- Claims to be the only such library: The authors state that existing non-autoregressive-capable libraries such as FairSeq and AllenNLP are no longer actively maintained and do not support non-autoregressive language modeling; to the authors' knowledge, XLM is the only library supporting fast prototyping of small non-autoregressive language models.
- Design principle of maximal independence: Core components delegate model-specific logic to pluggable instances, so the
Harnesscarries aModel,LossFunction, andPredictor, while theDataModulecarries oneDatasetManagerper dataset. Any of the four components can be swapped for a compatible alternative without changing the others. - Arbitrary component swapping without code changes: Hydra's
hydra.utils.instantiatelets users replace entire components purely from configuration files, enabling modular, nested configs and runtime code injection. - Reproduction of planning results: On synthetic path finding on star graphs, reproduced results (Table 1) are reported to be within 2% of the original papers' reported numbers.
- Planning task accuracies (sequence / token): ARLM 33.1 / 81.7 on Easy, 77.2 / 82.1 on Medium, 25.2 / 43.7 on Hard; ILM 100.0 / 100.0 on Easy, 100.0 / 100.0 on Medium, 97.5 / 98.2 on Hard; MLM 100.0 / 100.0 on Easy, 83.1 / 98.0 on Medium, 25.3 / 79.6 on Hard; MDLM 100.0 / 100.0 on Easy, 36.5 / 90.6 on Medium, 21.0 / 54.9 on Hard.
- Language modeling results on LM1B (Table 2): A 12-layer transformer was trained as ARLM, MDLM, and ILM respectively, with negative log-likelihood under Llama 3.2 8B used as the metric. NLL values are reported as 3.71 for corpus, 3.94 for ARLM, 4.81 for MDLM, and 4.72 for ILM. Entropy values are reported as 3.08 for corpus, 3.12 for ARLM, 3.70 for MDLM, and 2.81 for ILM. The authors state these are close to the results reported in Patel et al. (2025) for the same settings.
- Three main workflows:
train,eval, andgenerateare run via a single command withjob_type,job_name, andexperimentarguments, plus adebug=overfitsetting that overfits on a single batch for quick debugging. - Included building blocks: The library supplies modules including a standard decoder-only transformer, Diffusion Transformers, a rotary embedder, a time embedder, adaptive layer normalization layers, and standard noise schedulers; it also provides preconfigured datasets (StarEasy, StarMedium, StarHard, LM1B, OpenWebText), preconfigured models (ARLM, MDLM, ILM), loggers (TensorBoard default, WandB), metrics (accumulated loss, exact match, token accuracy), prediction logging,
push_to_hub, and callbacks such asEMACallback,ModelCheckpoint,SpeedMonitorCallback,ThinningCheckpoint, andOnExceptionCheckpoint.
Methodology in Plain English
The authors did not propose a new language model. Instead, they built infrastructure and then used it to re-run known experiments.
The framework separates concerns into two groups. Core components handle the execution flow that is the same for every model: a Harness (a PyTorch Lightning module) that instantiates the model, loss, predictor, and metrics and delegates to them; and a TextDataModule that manages an arbitrary number of datasets through one DatasetManager per dataset, each owning downloading, preprocessing, caching, a collator, and data-loader options. Model-specific components — the neural network, the loss function, the predictor (inference loop), and the collators — live in the model's own folder and are configured through Hydra files rather than code edits.
To show the framework works, the authors walk through implementing the Insertion Language Model step by step on a synthetic seq2seq star-graph path-finding task, then show the same steps for unconditional language modeling on LM1B. They then train ARLM, MDLM, MLM, and ILM with hyperparameters taken from prior work and compare against the original reported numbers, using negative log-likelihood under Llama 3.2 8B as the LM1B metric so the models are comparable.
Why This Matters
Impact on research: The main obstacle the paper targets is comparability. When every non-autoregressive method ships its own collation, loss, and decoding logic, differences in benchmark scores can come from implementation details rather than the modeling idea. A shared, model-independent core plus per-model pluggable components lets researchers isolate what actually changed, and the deliberate "copy over branching" principle keeps each model folder simple enough to share or lift out of the framework entirely.
Real-world applications (as framed or implied by the paper):
- Faster text generation, since non-autoregressive generation is pursued for its potential for faster inference speeds and better generation quality on certain tasks.
- Structured reasoning and planning tasks, such as generating a path through a graph given an edge list and start/goal nodes, which the paper uses as its representative seq2seq benchmark.
- Text generation via alternative paradigms, including masked diffusion, Gaussian diffusion, insertion, and edit-based models, each of which has different decoding requirements.
- Molecule generation and other non-text sequence generation, which the authors list as a planned extension.
Industry relevance: The package targets interoperability with existing infrastructure rather than replacing it — it inherits PyTorch Lightning features such as logging, checkpointing, and saving, uses standard HuggingFace-style datasets, and provides a push-to-hub command that reconstructs the training environment (datamodule, tokenizer, model architecture) from Hydra configs for upload to the Hugging Face Hub. That makes it practical for teams that want to prototype and package small models without building an evaluation harness from scratch.
Future Directions
- Add reference implementations and benchmark more models, including newer approaches such as Havasi et al. (2025) and Kim et al. (2025), which the conclusion and Appendix H name explicitly.
- Extend to non-text sequence generation, such as molecule generation and path planning, by adding support for external non-text datasets.
- Integrate FlexAttention, introduced in PyTorch 2.5, to enable fast attention with arbitrary masks and sequence packing, potentially eliminating the need for padding even for non-autoregressive models.
- Grow the suite of small pre-trained models distributed through the companion
xlm-modelspackage so the research community has reusable checkpoints, with further benchmarking work described as forthcoming.
Target Audience
Researchers and engineers who prototype, benchmark, or compare text-generation methods other than left-to-right decoding — particularly those working on masked diffusion, Gaussian diffusion, insertion, or edit-based language models. It is also useful for graduate students and practitioners who want a working, configurable codebase for small non-autoregressive models without writing collation, loss, and decoding infrastructure from scratch. Readers looking for new modeling techniques, new architectures, or state-of-the-art results will not find them here; the contribution is tooling, reproducibility, and workflow.
Authors’ abstract
In recent years, there has been a resurgence of interest in non-autoregressive text generation in the context of general language modeling. Unlike the well-established autoregressive language modeling paradigm, which has a plethora of standard training and inference libraries, implementations of non-autoregressive language modeling have largely been bespoke making it difficult to perform systematic comparisons of different methods. Moreover, each non-autoregressive language model typically requires it own data collation, loss, and prediction logic, making it challenging to reuse common components. In this work, we present the XLM python package, which is designed to make implementing small non-autoregressive language models faster with a secondary goal of providing a suite of small pre-trained models (through a companion xlm-models package) that can be used by the research community. The code is available at https://github.com/dhruvdcoder/xlm-core.