Skip to content
AI.info

Research

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation Overview Research area: Computer Vision — discrete video tokenization and autoregressive (AR) video generation / vid

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
arXiv
2610.00686
Published
2026-09-30
Authors
Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss

AI summary

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Overview

Research area: Computer Vision — discrete video tokenization and autoregressive (AR) video generation / video world models.

Technical level: Advanced. The paper assumes familiarity with VQ/FSQ tokenizers, representation alignment (REPA), flow-matching diffusion decoders, nested dropout, and autoregressive sequence modeling.

Scope in one sentence: The paper introduces SemanTok, a flexible-length, coarse-to-fine video tokenizer that injects frozen DINOv2 features into its encoder and trains lightweight heads to reconstruct those features from each retained token prefix alone, and it compares this against the closely related VideoFlexTok tokenizer across seven AR model sizes, two datasets, and a sweep of token budgets.

What This Paper Is About

Flexible-length video tokenizers promise that an autoregressive model can stop at any prefix length k and still decode a usable video, but existing tokenizers (notably VideoFlexTok) supervise semantics only through a REPA loss on an early decoder layer — a layer that also receives the noised latent and can partly satisfy that target without help from the tokens. The goal of SemanTok is to force the tokens themselves to carry the clip's semantics early, so that a small, cheap AR model can settle what the video contains before later tokens (or a larger model) spend capacity on pixel detail. The paper asks how this change in supervision affects generation fidelity, semantic alignment, generalization to unseen classes, and the cost of predicting tokens.

Key Contributions

  1. SemanTok tokenizer. A flexible video tokenizer that keeps the VideoFlexTok recipe (codebook, nested dropout, time-causal decoder, decoder REPA loss) but feeds frozen DINOv2-L patch features into the encoder alongside the VAE latents and adds a zero-initialized projection of the matching DINO class token into the first register token.

  2. Prefix-only semantic supervision. Two independently parameterized cross-attention heads — a Dense DINO head (spatial readout queries reconstructing DINO patches) and a Class DINO head (a single readout query for the class token) — read only the kept token prefix and never see the noised latent, so only the tokens can lower their losses. The full objective is L = L_base + 0.5·L_dense + 0.5·L_cls.

  3. A scaling and efficiency study. Seven AR sizes (49M to 2.29B parameters), plus a 1.33B AR model trained from 1.3B to 65.5B tokens, evaluated across token budgets k ∈ {1, 2, 4, …, 256} on Kinetics-600 (class-to-video) and uCO3D (text-to-video), reporting fidelity (gFVD, gFID), semantic alignment (class accuracy, ViCLIP, ClipV), tokenizer reconstruction (PSNR, SSIM, rFVD), and bits per token.

  4. A characterization of where the gain comes from — and where it fails. Evidence that SemanTok's generation advantage is concentrated in the first tokens and that its weakness appears at k = 1.

Main Findings

  • Small SemanTok models match much larger VideoFlexTok models. A 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size. Every SemanTok AR model reaches lower gFVD and gFID and higher class accuracy than the same-size VideoFlexTok model on both datasets. On Kinetics-600, SemanTok lowers gFVD by 11–24% and gFID by 5–17% and raises class accuracy by 25–61% (largest gains at the smallest AR sizes); on uCO3D the improvements are 2–13% (gFVD), 5–8% (gFID), and 22–30% (class accuracy). At k = 16, an 85M SemanTok AR model beats VideoFlexTok AR models of every size and budget up to 2.29B in gFID on both datasets and in gFVD on Kinetics-600. Even the smallest SemanTok model (49M) beats the best VideoFlexTok model up to 2.29B in class accuracy, ClipV, and ViCLIP on both datasets; on Kinetics-600, SemanTok's class accuracy is 0.631 versus 0.560 for the 2.29B VideoFlexTok model.

  • Scaling AR size improves fidelity more than semantic alignment. From 49M to 2.29B, SemanTok's best-k gFVD falls from 224 to 202 on Kinetics-600 and from 218 to 197 on uCO3D, and SemanTok's gFID advantage over VideoFlexTok at k = 64 halves from 49M to 2.29B. For the 1.33B AR model on Kinetics-600, SemanTok's best-k gFVD lead shrinks from 28% to 9% by 26B tokens and then holds at 11–14%; its class-accuracy lead persists at 30% after 65.5B tokens, so 5× more training buys VideoFlexTok fidelity but not semantic alignment. On uCO3D, SemanTok's ClipV nearly saturates by 13B tokens.

  • Semantic alignment generalizes to unseen classes. SemanTok prefixes recover object class at lower k than VideoFlexTok's on both in-distribution and out-of-distribution (OOD) uCO3D clips. With a 201M AR model at k = 16, SemanTok raises generated class accuracy by 24% on ID clips and 29% on OOD classes, and improves ClipV on both splits (Table 1). At k = 16, SemanTok reconstructions score higher ClipV and class accuracy than VideoFlexTok's but about 2.7 dB lower PSNR. ViCLIP favors SemanTok from k = 16 on, but not at k = 4.

  • The decoder's DINO semantics depend far less on the noised latent. At every k from pure noise, SemanTok's decoder-REPA readout has higher DINOv2 cosine similarity than VideoFlexTok's. At k = 32 on uCO3D, SemanTok's pure-noise readout reaches a similarity that VideoFlexTok reaches only with a 75%-clean latent (0.721 vs. 0.716). At k = 256, the latent adds almost nothing to SemanTok's readout but over 11× more to VideoFlexTok's. Appendix C reports that at σ = 0.25 VideoFlexTok's readout barely depends on k (0.716 at k = 1 to 0.719 at k = 256 on uCO3D; 0.674 to 0.685 on Kinetics-600).

  • Smaller realization gap. VideoFlexTok reconstructs somewhat better in PSNR and rFVD, but loses more fidelity in generation, so its gap between reconstruction and generation is wider. SemanTok has lower gFVD at every AR size at both k = 64 and k = 128. (On Kinetics-600, generated class accuracy can exceed reconstruction because the AR model sees the class label.)

  • The generation gain is concentrated in early tokens, which are cheaper to predict. At 201M with nothing forced, SemanTok's gFVD is 23% lower than VideoFlexTok's; forcing the first 16–64 ground-truth tokens per frame removes that lead. At k = 16, SemanTok needs 32% fewer bits per token (8.8 vs. 12.9 at 201M); since its marginal entropy is only about one bit lower, most of the saving comes from context, not from repetition. SemanTok's first 4 tokens cost 37 bits per frame yet nearly match the class accuracy of VideoFlexTok's first 32 tokens (405 bits). VideoFlexTok's gFVD degrades after only k = 16, while SemanTok's does not degrade until after k = 32.

  • Exception at one token per frame. At k = 1, SemanTok trails VideoFlexTok in semantic alignment with a mostly clean latent. All of SemanTok's losses that clear the bootstrap intervals are on uCO3D at k ≤ 4, mostly gFID and ClipV at k = 1; on Kinetics-600 SemanTok is never significantly worse. At k = 1 its pure-noise decoder-REPA lead is only 0.01–0.02 DINOv2 cosine versus 0.04–0.08 from k = 16, and with σ = 0.25 its readout is lower up to k = 4 on uCO3D and k = 8 on Kinetics-600. A single token carries at most about 16 bits — too few to satisfy flow matching, decoder REPA, and semantic supervision simultaneously. SemanTok nevertheless never falls behind in class accuracy at any budget or AR size. The stated limitation is that prioritizing semantic alignment over reconstruction makes it harder to reproduce colors and appearance details at low token budgets.

Methodology in Plain English

The researchers start from VideoFlexTok, which turns a clip into a sequence of register tokens using a time-causal encoder, quantizes them with FSQ (six dimensions rounded on an [8,8,8,5,5,5] lattice, giving a 64,000-entry codebook), applies nested dropout so any prefix stays decodable, and trains a flow-matching decoder jointly with the encoder. Concretely, a frozen VidTok VAE produces latents with T = 5, h = w = C_z = 16 (P = 256 patches per frame), features are lifted to width d_e = 1152, and K = 256 register tokens per frame are packed after the patches. Nested dropout samples k uniformly from {1, 2, 4, …, 256} per clip.

SemanTok changes two things. First, the encoder's inputs: each VAE patch is concatenated with the aligned frozen DINOv2-L patch feature before projection, and a zero-initialized projection of DINO's class token is added to the first register token. Second, the supervision target: two cross-attention heads (width d_h = 768, 12 attention heads, two layers each) read only the kept token prefix up to and including frame t, and try to reconstruct the DINO patches and the DINO class token by cosine loss. Because nested dropout varies k during training, every sampled prefix must support both predictions, so the most semantically relevant information is pushed into the earliest tokens. No loss assigns a specific DINO feature to a specific token, and the decoder's target remains the VAE latent, never DINO features.

For evaluation, both tokenizers use the same backbone, decoder REPA, codebook, sequence length, sampling settings, LLaMA-style AR recipe, and pixel resolution (17 frames at 128×128). All AR models train on all 256 tokens per latent frame; the budget k is chosen only at evaluation. Class accuracy on Kinetics-600 is closed-set UMT-L top-1 over 2,048 generated clips; on uCO3D it is nearest-class-mean top-1 in InceptionV3 space over 2,560 clips per split. Decoder-REPA probing uses 256 validation clips per dataset and tokenizers trained on 66B tokens (uCO3D) and 131B tokens (Kinetics-600); at k = 256 and σ = 0.25 the readout lies within 0.05 of the alignment each trainer logged for the same checkpoint.

Why This Matters

Impact on research. The paper argues that semantic supervision placed on an early decoder layer is partly satisfied by the noised latent rather than by the tokens, and that moving the target to prefix-only prediction heads changes what the discrete code actually contains. It reframes flexible tokenizers as a mechanism for deciding where in the token sequence semantics versus pixel detail live, connecting tokenizer design to the compression–generation trade-off and to work like FlexTok, VideoFlexTok, ReToK, LoST, and LARP.

Real-world applications (potential, based on the paper's framing):

  • Budget-aware video generation: an application can pick k at inference time and pay proportionally less compute, which the paper shows is now more favorable for SemanTok over most of the compute range (its envelope of best score versus AR FLOPs is better).
  • Text-to-video over objects (uCO3D) and class-conditioned action video (Kinetics-600), including categories never seen in training.
  • Compact state representations for downstream models, which the authors list as a future direction for action-conditioned world models.
  • Content creation pipelines where a coarse prefix fixes scene content cheaply and later tokens refine appearance and motion.

Industry relevance. The affiliation includes Stability AI, and the central claim is a practical cost argument: a 201M SemanTok AR model matching or beating a VideoFlexTok model 3.4× its size, and an 85M model beating VideoFlexTok models up to 2.29B at k = 16, implies substantially cheaper inference and training for a given quality bar. The 32% reduction in bits per token at k = 16 under the same 201M AR model is a direct prediction-cost saving.

Future Directions

  • Optimize predictability directly, rather than relying on semantic supervision as a proxy for tokens that are both easy to predict and informative.
  • Use video-native or language-aligned teachers to organize the early prefixes, extending the current frozen DINOv2-L supervision.
  • Use semantic prefixes as compact states for action-conditioned world models, where an early, semantically fixed prefix could serve as a world-state representation.
  • Address the k = 1 failure mode: a single token carries at most about 16 bits, which is insufficient for flow matching, decoder REPA, and semantic supervision simultaneously, and SemanTok lags VideoFlexTok in semantic alignment at k ≤ 4 on uCO3D — an open question is whether the very shortest budgets can be served without trading off reconstruction of colors and appearance.

Target Audience

Researchers and engineers working on video tokenization, autoregressive video generation, and video world models; practitioners who need to trade generation compute against quality at inference time; and readers interested in how representation-alignment losses interact with the noised-latent shortcut in diffusion decoders. The paper is most useful to those already comfortable with FSQ/VQ codebooks, REPA-style alignment, and flow-matching decoders, since Appendix D and the reproducibility statement assume that background.

Authors’ abstract

Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.

Read the original paper