Research
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
Overview Research area: Natural language processing — specifically tokenization, vocabulary design, and the geometric structure of word embeddings in large language models (LLMs). Technical level: Int

- arXiv
- 2510.17001
- Published
- 2025-10-19
- Authors
- Yuval Reif, Guy Kaplan, Roy Schwartz
AI summary
Overview
Research area: Natural language processing — specifically tokenization, vocabulary design, and the geometric structure of word embeddings in large language models (LLMs).
Technical level: Intermediate. The paper is readable without deep background, but it assumes familiarity with tokenization (BPE), embedding and unembedding matrices, and lightweight fine-tuning methods such as LoRA.
Scope: The paper shows that word-form variation (e.g., walk vs. walked) is encoded as a linear direction in LLM embedding space, and uses that property to replace unique tokens for surface forms with compositions of shared "base form" and "transformation" vectors, freeing 10–40% of vocabulary slots while maintaining downstream performance.
What This Paper Is About
Standard tokenizers like byte-pair encoding assign separate vocabulary entries to every surface form of a word — walk, walks, walking, walked — so a size-capped vocabulary fills up with redundant form variation at the expense of rare words and multilingual coverage. The authors test whether LLMs already represent these relationships as simple linear offsets, and if so, whether whole surface forms can be reconstructed by adding a learned "transformation" vector to a shared "base form" embedding. The goal is a vocabulary that is simultaneously more compact and able to cover more words, without retraining the model backbone.
Key Contributions
-
A quantification of redundancy in existing vocabularies. Analyzing the English whole-word tokens of the GPT-4 tokenizer against UniMorph's English lexicon, the authors identify 24.6k such tokens. Ignoring case reduces this to 17.7k unique types, and further collapsing inflectional and derivational relations leaves just 14.3k base forms — a total reduction of 42%.
-
Transformation vectors as an input and output mechanism. The paper formalizes a compositional vocabulary in which a word w is represented as a base form plus a set of transformation vectors (Eq. 1), and replaces the model's single unembedding matrix with separate base-form and transformation matrices whose dot-products are summed to produce logits (Eq. 2). This applies to both in-vocabulary and out-of-vocabulary words, as long as the base form is in-vocabulary.
-
Evidence that pretrained LLMs already interpret compositions correctly. Using Patchscopes probing across five languages and five models, the authors show that inflectional and capitalization transformations are often resolved as the intended word, sometimes even for words that never existed as single tokens in the vocabulary.
-
Two practical regimes with measured vocabulary savings. Post-hoc adaptation of pretrained models (frozen backbone, LoRA on the final k = 8 blocks, fewer than 0.001% additional parameters) removes up to 10% of the vocabulary; compositional pretraining from scratch removes roughly 41% of BPE vocabulary entries.
Main Findings
-
Inflection and capitalization compose; derivation does not. In Llama-3.1-8B English experiments, plural nouns composed from base + transformation were resolved correctly 92% of the time at the embedding layer and 96% after early-layer detokenization, and capitalization 80%/89%. Derivational suffixes fared far worse: the average over all derivatives was 31% at the embedding layer and 45% after detokenization. The authors suggest derivations rarely occur as in-vocabulary single tokens, so models learn weaker linear structure for them.
-
Composition works for many out-of-vocabulary words. For out-of-vocabulary English plurals the embedding-layer accuracy was 30% and detokenization accuracy 56% (N = 3.4k); present singular verbs reached 64%/82%; capitalization reached 72%/85% (N = 8.4k). Out-of-vocabulary derivations were almost never resolved (0%/3% across 31.4k evaluated forms).
-
Compositional interpretation generalizes across languages and models. The paper reports results for ALLaM-7B on Arabic, EuroLLM-9B on German, Russian and Spanish, and Llama-3 on English. Examples: Spanish capitalization 97% in-vocabulary and 90% out-of-vocabulary; Russian capitalization 97% in-vocabulary; ALLaM Arabic noun inflection 77% in-vocabulary and 14% out-of-vocabulary.
-
Some transformations generalize better outside English than inside it. The authors note that adjective and verb inflections sometimes work better for out-of-vocabulary representation in other languages than in English, which they connect to the observation in Section 8 that models with smaller token vocabularies learn stronger linear morphological encodings.
-
Post-hoc adaptation costs little. For Llama-3.1-8B English across MMLU, ARC, BoolQ, TriviaQA, SQuAD, HellaSwag, Winogrande, PIQA and COPA, the average score moved from 66.9 to 65.9, a change of −1.0. Individual deltas ranged from +0.5 (Winogrande, 78.1 to 78.6) to −3.3 (TriviaQA, 66.5 to 63.3). Multilingual deltas across XNLI, XQuAD and Global MMLU ranged from +0.6 (EuroLLM German XNLI, 46.5) to −4.5 (EuroLLM Russian XNLI, 40.1).
-
Vocabulary slots freed. The approach removes roughly 10k surface-form tokens from Llama3 and OLMo2 each, and 7.8k from Qwen2.5. In the other four languages, absolute reductions are smaller (0.6k–3k) but represent 38–45% of whole-word tokens in the target language. Decoding speed drops by 0.8%.
-
Reallocating freed slots improves tokenization efficiency. Simulating eviction of 10k English surface words from the Llama-3.1-8B tokenizer and adding 2.5k language-specific BPE tokens for each of Arabic, Russian, German and Spanish improved bytes-per-token on held-out FineWeb-2 text from 4.40 to 4.81 on average (+9.3%), with Arabic improving from 4.62 to 5.46.
-
Compositional vocabularies work when training from scratch. Reshaping a 50k-token GPT-2 tokenizer removed 41.6% of vocabulary entries with bits-per-byte moving from 1.08 to 1.09. Reshaping a 32k-token Spanish BPE vocabulary removed 41.8% of entries, with BPB moving from 1.00 to 1.11 and bytes-per-token improving from 4.77 to 4.92 when the model was allowed to generate out-of-vocabulary surface forms.
-
Morphological linearity weakens as vocabulary grows. Across models with English vocabulary sizes spanning 8k–44k tokens and total vocabularies of 32k–256k tokens, models with compact English vocabularies (8–10k words, e.g., Llama2, Mistral) encode morphology as generalizable vector offsets, while large-vocabulary models (~40k words, e.g., Falcon3, Gemma2-9B) tend to treat inflected forms as individual lexical units, with weight tying amplifying the trend.
-
Note on the extracted text. Some cell values in the multilingual results table are partly garbled in the available text (for example, several Russian and German sample sizes and one German adjective-inflection figure). Values that are not clearly legible are omitted here rather than guessed.
Methodology in Plain English
The authors first look for structure rather than training anything. They take a model's existing vocabulary and ask: which tokens are just morphological variants of other tokens? They use UniMorph, a multilingual database of word forms, to map each surface form to a base word plus a set of standard morphological tags (such as V;PST for past tense), adding their own rules for capitalization.
They then compute a transformation vector for each relation by taking the average difference between the embeddings of surface forms and their base forms, summed only over words that exhibit a single transformation, so the signal stays clean. The same calculation is done separately in the input embedding space and the output unembedding space.
To check whether a pretrained model actually understands these composed vectors, they feed a composed embedding — base plus transformation — into the model and use Patchscopes, a prompting technique that asks the model to describe the contents of a hidden state in natural language. If the description matches the intended surface form, the composition "worked." They check this at the embedding layer and across the first ten layers (detokenization).
For end-to-end language modeling, they modify the model so that input embeddings are looked up compositionally and next-token logits are computed by summing base-form and transformation scores before a single softmax. Only the transformation vectors and LoRA adapters on the final eight blocks are trained; everything else stays frozen. They train with two-stage knowledge distillation on roughly 5M tokens of FineWeb-Edu at sequence length 256, first fitting input transformations against the original model's predictions and then output transformations against the resulting model. Compositions that failed the earlier evaluation, plus all derivational transformations, are filtered out and fall back to the original tokenization.
Finally, they test the alternative regime: pretraining small nanoGPT-124M models on 1B tokens with compositional vocabularies from scratch, where the next token is predicted by first sampling a base form and then a transformation conditioned on it, and performance is measured in bits-per-byte so it is comparable across different vocabularies.
Why This Matters
Impact on research. The paper reframes vocabulary design as a
Authors’ abstract
Large language models (LLMs) often encode word-form variation (e.g., walk vs. walked) as linear directions in the embedding space. However, standard tokenization algorithms treat such variants as distinct words with different vocabulary entries, quickly filling the size-capped token vocabulary with surface-form variation (e.g., walk, walking, Walk) at the expense of diversity and multilingual coverage. We show that many of these variations can be captured by transformation vectors: additive offsets that yield the appropriate word representation when applied to a base form embedding, in both the input and output spaces. Building on this, we propose a compact reshaping of the vocabulary: instead of assigning unique tokens to each surface form, we compose them from shared base form and transformation vectors (e.g., walked is walk+past tense). Our approach is lightweight, keeping the pretrained backbone frozen and only training small adaptation modules. We apply it across five languages and multiple LLMs in both pretraining and post-hoc adaptation, freeing 10-40% of vocabulary slots to be reallocated where tokenization is inefficient. Importantly, we do so while also expanding vocabulary coverage to out-of-vocabulary words, and with minimal impact on downstream performance. Our findings motivate a rethinking of vocabulary design, towards a representation that better matches the underlying structure of language and the practical needs of multilingual coverage.