Skip to content
AI.info

Research

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Overview Research area: Natural Language Processing — specifically, the architecture and parameterization of transformer language-model inputs, at the boundary between tokenization and learned represe

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
arXiv
2610.04002
Published
2026-10-02
Authors
A. Bochkov

AI summary

Overview

Research area: Natural Language Processing — specifically, the architecture and parameterization of transformer language-model inputs, at the boundary between tokenization and learned representation.

Technical level: Advanced. The paper is short and empirically focused, but it relies on concepts such as injective encoding, linear algebra over GF(2), parameter accounting, and standard base-model benchmark suites.

One-sentence scope: A controlled, backbone-matched comparison of a learned input embedding table against two fixed token-identity interfaces (canonical 16-bit binary codes and one invertible GF(2) recoding) in three 1.7B-class decoder-only language models trained from scratch on a 100B prediction-token budget.

What This Paper Is About

Nearly every neural language model gives each vocabulary item its own independently trainable input vector, and that row is optimized along with the rest of the network. This paper asks whether that specific parameterization is actually required for substantial language-modeling capability, or whether a shared transformer can learn useful language computation from token identities that are fixed and never updated. The author trains three otherwise identical models that differ only in their input interface, and measures how much capability survives when the trainable input table is removed entirely.

Key Contributions

  1. A backbone-matched comparison of learned versus fixed token-identity interfaces at 1.7B-class scale, using one standard tokenizer family, a shared contextual architecture, an untied output head, and a broad base-model evaluation suite.
  2. A direct measurement of both the capability retained and the quality cost incurred when all independently trainable token-specific input vectors are removed — reported as a trade-off rather than a win.
  3. A noncanonical fixed-code experiment using one invertible linear recoding over GF(2), showing viability under a changed token-to-code assignment, while explicitly not claiming invariance to code assignment in general.
  4. An explicit account of what an immutable input interface does and does not control, offered as a research instrument for studying representation learning downstream of an input whose coordinates never move.

Main Findings

  • Fixed codes support substantial capability. The canonical binary-code model reaches 52.40% HellaSwag normalized accuracy, 70.51% PIQA accuracy, and 42.75% LAMBADA accuracy. The GF(2) recoded model reaches 51.44% HellaSwag normalized accuracy, 71.16% PIQA accuracy, and 42.29% LAMBADA accuracy. Both also achieve useful corpus likelihood on WikiText.

  • The learned-input control still wins on several evaluations. The learned-table model scores 57.79% HellaSwag normalized accuracy, 66.04% ARC-Easy normalized accuracy, 37.63% ARC-Challenge normalized accuracy, 72.69% PIQA accuracy, 58.56% WinoGrande accuracy, 37.80% OpenBookQA normalized accuracy, and 47.72% LAMBADA accuracy. The paper therefore states explicitly that this establishes viability rather than performance parity.

  • The recoded assignment is viable but not shown to be equivalent. The GF(2) model performs in the same broad range as the canonical model: canonical codes are better on HellaSwag and WikiText (18.58 versus 19.03 word perplexity), while the ordering reverses on some other metrics. Only one recoding matrix was evaluated, generated once with code seed 12345.

  • Near-chance tasks bound the claim. CommonsenseQA accuracy is approximately 20% for all three models, and MMLU stays approximately 25–26% at both zero and five shots. The author reports these rather than reinterpreting them as evidence of broad knowledge or reasoning.

  • External references calibrate absolute quality. Canonical-code HellaSwag normalized accuracy of 52.40% sits between SmolLM2-135M at 43.02% and SmolLM2-360M at 56.28%; SmolLM2-1.7B is substantially stronger (71.43% HellaSwag normalized accuracy). SmolLM2's reported training budgets are approximately 2T, 4T, and 11T tokens respectively.

  • The parameter reduction is real but secondary. The fixed interfaces remove 100,663,296 trainable input parameters (approximately 5.56% of the untied learned-input model), yielding 1.711B-parameter models (1,711,376,384) versus 1,812,039,680 for the learned control. The author states that parameter reduction is not the central result.

  • Identity, learnability, and utility are distinct questions. The paper separates whether the input preserves which token was observed (identity), whether a shared network can learn from that interface (learnability), and whether per-token independent adaptation improves the model (utility). Only the third is answered in the negative.

  • No systems-efficiency advantage is claimed. The paper reports no measured throughput, energy, or latency advantage, and notes that removing the input use of a matrix need not remove the matrix from storage when input and output weights are tied.

Methodology in Plain English

The author trains three small-but-serious language models from scratch, roughly 1.7–1.8 billion parameters each, on a target budget of 100 billion prediction tokens per model. Everything is held constant except the very first step: how a token ID becomes a vector that enters the transformer.

  • Model 1 (control): a conventional learned input embedding table, where each of the 49,152 vocabulary items has its own adjustable vector.
  • Model 2 (Binary16): each token ID is written as its 16-bit little-endian binary expansion. Since 2^16 is larger than 49,152, 16 bits are enough to give every token a distinct pattern. Those 16 zeros-and-ones are simply repeated (tiled) 128 times to reach the model's width of 2048, with no trainable projection in between. No code vector is learnable.
  • Model 3 (GF2): the same idea, but the 16-bit codes are first scrambled by a fixed invertible matrix over GF(2), with an offset of zero. Because the matrix is invertible, token identity is still preserved; the token-to-bit-pattern assignment changes.

All three use the same tokenizer, 24 decoder blocks, hidden width 2048, 32 attention heads, RoPE, RMSNorm, SwiGLU with intermediate width 8192, context length 2048, and an untied trainable output head. Training uses AdamW with a peak learning rate of 1.5e-4, 2000 warmup updates, cosine decay, gradient clipping at 1.0, on two GPUs with 8 sequences per GPU and 8 gradient-accumulation steps — 262,144 prediction targets per optimizer update.

Evaluation uses the EleutherAI Language Model Evaluation Harness on HellaSwag, ARC, PIQA, WinoGrande, OpenBookQA, CommonsenseQA, MMLU (zero-shot and five-shot), LAMBADA, and WikiText. The three SmolLM2 base checkpoints are scored with the same local protocol purely as external quality references, not as trained controls.

The author is explicit that this is one training run per interface, that evaluation standard errors are not training-seed variability, that input trainability, parameter count, geometry, and input scale all change together, and that the quality gap therefore cannot be attributed to a single mechanism.

Why This Matters

Impact on research. The paper reframes a widely assumed architectural necessity as an empirical choice. A trainable input table is shown to be useful but not required for the capabilities observed at this scale. Equally important is the methodological point: separating "does this component help" from "is this component required" prevents a common conflation in architecture research. The fixed-input setting also provides a controlled boundary condition — when input coordinates never move, any change in behavior cannot be attributed to them, which is useful for studying how representations are built.

Real-world applications (as the paper's framing supports them):

  • Embedding-free or reduced-storage input interfaces for models where vocabulary-sized input matrices dominate parameter counts, particularly in multilingual or very large-vocabulary settings.
  • Compact on-device or embedded deployment, where removing roughly 100.7M trainable input parameters (and their optimizer state) is attractive — though the paper reports no measured latency, throughput, or energy benefit.
  • Experimental infrastructure for interpretability and representation research, where an immutable input makes it easier to track how lexical and contextual distinctions emerge across training and depth.
  • Modular system design, where the token-identity interface stays fixed while contextual computation or memory components are swapped, retrained, or studied independently.

Industry relevance. The result is relevant to anyone making parameter-allocation decisions at the vocabulary boundary. The paper notes that weight tying and factorized embeddings already show this boundary is an architectural choice, and it explicitly cautions that its parameter savings are measured against an untied baseline, so they should not be read as savings over a tied model. The headline for practitioners is not "delete your embedding table" but "the table is a design lever, and here is what removing it costs."

Future Directions

  • Separate fixedness from geometry. The paper proposes comparing learned low-dimensional coordinates, fixed random continuous codes, norm-matched controls, and multiple independently sampled recodings to determine whether the observed cost comes from fixedness, dimensional restriction, input scale, or the specific code assignment.
  • Study representation formation on a stationary input. Tracking how lexical and contextual distinctions become accessible across training and depth when raw token coordinates never change — paired with causal interventions such as layer ablations and controlled activation substitutions rather than clustering evidence alone.
  • Test generalization beyond the tested configuration. The experiment uses one tokenizer family, one vocabulary size, one context length, one data configuration, and one model scale; multilingual behavior, long-context operation, instruction tuning, and substantially different budgets remain untested.
  • Find a systems-level rationale or rule it out. Whether removing the input table yields any measured throughput, energy, or latency advantage is explicitly left open, since runtime depends on buffer access, tiling, activations, and kernel behavior.
  • Investigate the zero-code boundary case. Canonical coding maps token ID zero to the all-zero vector, and a bias-free architecture would produce zero hidden states for an all-zero input sequence. The paper notes this requires documenting how empty contexts and initial special tokens are handled during evaluation.

Target Audience

This paper is most valuable to language-model architecture researchers and graduate students who work on tokenization, embedding design, and parameter allocation — especially those interested in controlled ablation-style comparisons rather than headline benchmark gains. It also suits interpretability researchers looking for a setting where the input is immutable, and efficiency-minded practitioners who want a clear-eyed account of what is and is not gained by dropping a vocabulary-sized input matrix. Readers should be comfortable with transformer terminology and basic linear algebra; the paper is short and does not require deep mathematical background beyond that.

Authors’ abstract

A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.

Read the original paper