Research
UTF-8 Plumbing: Byte-level Tokenizers Unavoidably Enable LLMs to Generate Ill-formed UTF-8
Overview Research area: Natural Language Processing — specifically tokenization and the byte-level versus code-point-level design of large language model vocabularies, with a formal methods component.

- arXiv
- 2511.05578
- Published
- 2025-11-05
- Authors
- Preston Firestone, Shubham Ugare, Gagandeep Singh, Sasa Misailovic
AI summary
Overview
Research area: Natural Language Processing — specifically tokenization and the byte-level versus code-point-level design of large language model vocabularies, with a formal methods component.
Technical level: Advanced. The paper's central argument rests on a formalization of tokenization using monoid theory and a proof about sequences produced by byte-level tokenizers, which presupposes familiarity with tokenizer internals and formal language theory.
Scope: The paper proves that any tokenizer whose vocabulary contains tokens that are not well-formed UTF-8 can generate ill-formed UTF-8 output, shows a formal discrepancy between incremental and whole-sequence decoding, and examines mitigations and real-world case studies.
What This Paper Is About
Language models do not read raw text; they read sequences of vocabulary items ("tokens") produced by a tokenizer, and they generate output as sequences drawn from that same vocabulary. The vocabulary can be assembled out of code points, which guarantees every token is a valid UTF-8 character but demands thousands of initial vocabulary entries to cover input well, or out of raw bytes, which needs only 256 entries to avoid out-of-vocabulary errors but offers no such guarantee — neither individual tokens nor sequences of them are required to be valid UTF-8.
The paper's goal is to formalize this distinction and prove that the byte-level choice makes ill-formed UTF-8 output not merely possible but unavoidable, then to connect that formal result to real bugs in how systems and applications handle model-generated text.
Key Contributions
-
A formal theory of tokenization. The authors formalize tokenization using monoid theory, giving a mathematical setting in which the properties of token vocabularies and token sequences can be stated and proved.
-
A proof of unavoidable ill-formed output. They prove that any tokenizer whose vocabulary contains tokens that are ill-formed UTF-8 can always produce sequences that are ill-formed UTF-8 — that is, the problem is a structural consequence of the vocabulary design, not an occasional glitch.
-
A formal result on decoding order. They demonstrate formally that incrementally converting tokens back into a string and interpreting the intermediate results as UTF-8 yields different results from converting the entire sequence of tokens at once. This is a divergence between two decoding strategies that programmers might assume are equivalent.
-
Mitigations and case studies. The formal result is presented as a predictor of real-world bugs. The paper evaluates mitigations for the identified problem and provides case studies spanning major foundation models, serving engines, and constrained generation systems.
Main Findings
-
Byte-level vocabularies trade coverage for validity. Starting from bytes avoids out-of-vocabulary errors with only 256 initial vocabulary members, whereas a code-point vocabulary requires thousands of initial members to achieve acceptable coverage of inputs. The abstract frames this as the core trade-off driving the problem.
-
Ill-formed UTF-8 output is guaranteed, not incidental. The paper proves that tokenizers with ill-formed UTF-8 tokens in their vocabulary can always produce ill-formed UTF-8 sequences. This holds for the vocabulary members themselves and for sequences built from them.
-
Incremental and whole-sequence decoding disagree. The formal analysis shows that converting tokens to a string step by step and interpreting each intermediate result as UTF-8 gives different results than converting the whole token sequence at once — meaning the choice of decoding procedure changes the output, not just its performance.
-
Downstream code breaks. Sequences that are not valid UTF-8 break code that assumes its input is valid UTF-8. The abstract states that applications built on language models must account for the breakage this introduces.
-
Real systems are affected. The abstract claims case studies of major foundation models, serving engines, and constrained generation systems, and states that mitigations are evaluated. Specific systems, mitigation designs, and quantitative outcomes are not reported in the abstract.
Methodology in Plain English
The authors take an abstract mathematical view of what a tokenizer does. Tokenization is modeled as an operation that combines pieces of text, and the rules governing that combination are captured with monoid theory — a branch of mathematics concerned with how things are combined in sequence and what identities and associativity properties those combinations have. Within that framework, they ask what happens when the building blocks are bytes rather than complete characters, and they prove that ill-formed UTF-8 sequences must be reachable. They then compare two ways of reassembling a token sequence into text: piece by piece, checking UTF-8 validity as you go, versus all at once at the end. Having established the formal result, they treat it as a prediction about software behavior, evaluating proposed mitigations and examining how the issue plays out in real foundation models, serving engines, and constrained generation systems. The abstract does not describe the experimental setup, the mitigations themselves, or the metrics used to judge them.
Why This Matters
Impact on research. The paper reframes byte-level tokenization — widely adopted because it eliminates out-of-vocabulary errors cheaply — as carrying a provable correctness cost rather than a merely practical one. By grounding the argument in monoid theory, it gives the tokenization literature a formal vocabulary for a class of failure that has largely been treated as an engineering annoyance.
Real-world applications:
-
Language model serving and APIs. Serving engines that stream tokens to clients must decide how to buffer and validate output; the incremental-versus-whole-sequence divergence identified here sits directly in that path.
-
Downstream consumers of model output. Parsers, JSON and markup handling, database writes, and any code that assumes its input is valid UTF-8 can fail or silently misbehave on ill-formed sequences.
-
Constrained generation and structured output. Systems that restrict what a model may generate need to reason about validity at the token level, which is exactly where the paper's formal result applies.
-
Evaluation and benchmarking pipelines. Test harnesses and data-processing code that consume model output inherit the same assumption of valid UTF-8.
Industry relevance. Because byte-level tokenization is common precisely for its cheap coverage and lack of out-of-vocabulary failures, this is a problem that affects production systems by default rather than by accident. The abstract positions the findings as predicting real bugs, which makes the mitigations and case studies relevant to anyone operating or auditing model infrastructure.
Future Directions
-
Designing and standardizing mitigations. The abstract states that mitigations are evaluated but does not describe them; the natural next step is establishing which approaches are robust and how they should be specified for serving engines and client libraries.
-
Making decoding semantics explicit. Since incremental and whole-sequence decoding can differ, interfaces and specifications for token streaming need to state which behavior they guarantee and what consumers should expect.
-
Extending the formal framework. The monoid-theoretic treatment could be pressed further to characterize exactly which vocabularies can avoid ill-formed output, and at what cost in vocabulary size and coverage.
-
Broader case studies and tooling. The abstract cites case studies of foundation models, serving engines, and constrained generation systems; widening that coverage and turning it into diagnostics or validation tooling are open steps.
Target Audience
Researchers and engineers working on tokenization, language model serving infrastructure, and structured or constrained generation will find the core argument most directly useful. The formal methods framing makes it valuable to readers interested in applying verification techniques to machine learning systems. Practitioners who write or maintain code that consumes model output — parsers, API clients, data pipelines — benefit from the practical warning and the mitigation discussion. A background in tokenizer mechanics and some comfort with formal notation is needed to follow the proofs; the abstract alone does not supply them.
Authors’ abstract
Subword tokenization segments input text according to a pre-defined vocabulary to feed it into a language model; the language model, in turn, generates a sequence made from this same vocabulary. The members of the vocabulary can be built of code points or bytes. Using code points means that all members of the vocabulary are valid UTF-8 characters. However, it also requires thousands of initial members to achieve acceptable coverage of inputs. Beginning with bytes, on the contrary, avoids out-of-vocabulary errors with only 256 initial members of the vocabulary, but the members of the vocabulary and sequences of them are not guaranteed to be valid UTF-8. Sequences that are not valid UTF-8 break code that assumes its input to be valid UTF-8. Applications of language models must account for the breakage thereby introduced. In this paper, we formalize tokenization using monoid theory and prove that tokenizers whose vocabularies contain tokens that are ill-formed UTF-8 can always produce sequences that are ill-formed UTF-8. We demonstrate formally that attempting to incrementally convert tokens back to a string and interpret the results as UTF-8 gives different results than converting the whole sequence of tokens at once. This formal result predicts real-world bugs: we evaluate mitigations for the problem identified and provide case studies of major foundation models, serving engines, and constrained generation systems.