Research
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
Beyond Selection: Token Parameterization for Extreme Visual Token Compression Overview Research area: Computer vision / vision-language model efficiency — specifically visual-token compression for vis

- arXiv
- 2609.35232
- Published
- 2026-09-28
- Authors
- Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
AI summary
Beyond Selection: Token Parameterization for Extreme Visual Token CompressionOverview
Research area: Computer vision / vision-language model efficiency — specifically visual-token compression for vision-encoder → LLM pipelines.
Technical level: Advanced. The paper combines transform coding theory (DCT/Haar orthonormal bases), optimization diagnostics over Gram matrices and orthogonal re-parameterizations, and end-to-end multimodal benchmark evaluation on LLaVA-1.5-7B.
Scope: The paper reframes extreme visual-token compression as a token parameterization problem rather than a token-selection problem, proposes a four-step lightweight coder called Braco (Backbone–Residual + basis + coordinate), and evaluates it across 23×–144× compression against pruning, merging, learned-interface, and transform-coding baselines.
What This Paper Is About
Vision-language models spend cost proportional to the number of visual tokens they receive, but deployment settings such as mobile inference, low-latency interactive agents, and long-context multimodal reasoning may allow only dozens of tokens per image. Existing approaches — pruning or merging patches, or learning attention-based resamplers — either break down by discarding rare but critical regions under extreme budgets, or add attention computation, parameters, and training complexity. The paper asks whether performance under extreme token budgets is mainly a matter of selecting tokens more carefully, or of choosing a better parameterization for the visual token field, and answers with a lightweight coder that separates which subspace is retained from how coordinates inside that subspace are organized.
Key Contributions
-
A token-parameterization view of extreme visual-token compression. The paper separates basis transformation and structured truncation (which determine the retained subspace) from coordinate organization (which affects optimization and cross-modal alignment), rather than treating compression as token selection alone.
-
Unified compressibility and learnability diagnostics. Compressibility is formalized as a trade-off between energy retention and task-direction retention (a "readability" score), and learnability as a combination of a statistical-conditioning penalty and a geometric penalty on orthogonal coordinate changes within the same retained subspace.
-
Two compensation mechanisms for transform truncation. An input-independent basis-coordinate embedding restores stable token identity on the transform lattice (since DCT coefficients lose explicit spatial locality), and a lightweight sparse-pooling spatial residual module learns a small set of tokens to recover localized evidence lost under low-pass truncation.
-
A deployable four-step coder, Braco. The design combines transform-basis truncation, basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and a small spatial residual budget, and is evaluated end-to-end under matched token budgets.
Main Findings
-
DCT concentrates energy under deployable truncation. Under a fixed structured truncation rule, DCT and Haar retain far more token-field energy than spatial or random orthonormal bases — approximately 0.51 vs approximately 0.04 at K = 32, and approximately 0.57 vs approximately 0.08 at K = 64. The gap largely disappears under oracle magnitude truncation, showing that the fixed deployable ordering matters.
-
Basis choice also affects linearly readable task information. With frozen patch tokens and the same linear-probe setup on CelebA, DCT reaches 91.6% versus 89.8% for spatial at K = 1, and 92.5% versus 90.9% at K = 4.
-
Coordinate organization is budget-dependent. Holding the retained DCT subspace fixed and varying only the coordinates, time-to-threshold on held-out dev cross-entropy (< 2.50) shows that at K_b = 4 only vanilla coefficient coordinates reach the target, while at K_b = 16 and K_b = 64 the idct coarse-grid organization reaches it 18% and 60% faster. The implementation therefore uses vanilla when K_b = C² < 16 and idct when K_b ≥ 16.
-
Favorable accuracy–efficiency frontier at extreme budgets. Relative to the 576-token Vanilla upper bound, Braco reduces full-pipeline prefill FLOPs from 8.67T to 1.09–1.37T while retaining 91.2–95.2 Vanilla-normalized Acc. across 4–25 tokens. It attains 95.2% Acc. at 23×, 93.2–94.0% at 36×–64×, and 91.2% at 144×.
-
Leading accuracy at 25, 16, and 9 tokens, and near-parity at 4 tokens. Braco attains the highest Acc. among the evaluated methods at 25/16/9 tokens and remains within 0.2 Acc. of QueCC at 4 tokens with lower cost.
-
Much cheaper compressor module. At 16 tokens, Braco matches QueCC's accuracy with 16.6× lower module latency and 78.8× fewer compressor FLOPs. Measured pre-projector, Braco reports 94.0 Acc., 1.073 ms latency, 1.389 G FLOPs, and 8.731 MB memory, versus QueCC's 93.9 Acc., 17.864 ms, 109.504 G, and 14.875 MB.
-
End-to-end speedup against prior methods. Braco achieves up to approximately 36% end-to-end speedup; at 16 and 9 tokens it matches QueCC within 0.1 Acc. while reducing latency by about 36%.
-
Larger inputs favor the approach. Beyond the 576-token Vicuna-7B setting, Braco keeps 98.1 Acc. at 2880 input tokens with Vicuna-7B, reducing full-pipeline FLOPs from 40.57T to 3.45T and latency from 261.08 ms to 44.36 ms. With Qwen2.5-3B it retains 93.1/92.6 Acc. at 729/1024 input tokens and 90.8 Acc. at 3645 input tokens (182× compression), while QueCC degrades on Qwen2.5-3B. The abstract reports 90.8% Acc. at 182× versus 68.9% for a comparable prior.
-
Ablations support the hybrid design. At 9 tokens, the hybrid c2s5 allocation reaches 93.2 Acc., above pure-backbone c3s0 (89.8) and residual-only c0s9 (92.0) at comparable compressor cost. At the 16-token c3s7 setting, replacing the DCT backbone with spatial/Haar tokens costs 2.3–3.0 Acc. points, replacing the coordinate organization with idct/random rotation costs 2.5–3.1 points, and replacing Polar Fourier embeddings with learned or 2D sine–cosine variants costs 0.8–0.9 points.
Methodology in Plain English
The authors start by writing the compressed visual representation as a product of three choices: an orthonormal basis that re-expresses the patch-token grid, a fixed index set that keeps only a limited set of transform coefficients, and an orthogonal matrix that may reshuffle the retained coordinates without changing the information they carry. The first two choices decide which subspace survives compression; the third changes only the coordinates inside that subspace.
They then define two measurable objectives. Compressibility measures how much of the token field's energy a basis–truncation pair keeps, combined with how much of a downstream task direction that same retained subspace captures. Learnability measures how easy the retained coordinates are to optimize, using a penalty on off-diagonal and uneven diagonal entries of the coordinate Gram matrix, plus a penalty for drifting away from a preferred structured organization.
Applying these objectives, the authors select a 2D DCT as the basis and a C × C low-frequency block as the truncation rule, after comparing spatial, DCT, Haar, and random orthonormal bases. Because DCT coefficients are global rather than local, they add an input-independent embedding indexed by the retained transform coordinate to keep token identity stable. They then choose between coefficient coordinates and an inverse-DCT coarse grid according to the backbone size. Finally, because a spatially sparse residual is diffuse in an incoherent transform basis — bounded by m μ² s / L in the paper's equation — they add a small number of learned residual tokens produced by a TokenLearner-style scorer with sparsemax weights over the original spatial grid.
The total budget splits into a backbone of C² tokens and S residual tokens, instantiated as c1s3, c2s5, c3s7, and c4s9 for budgets K = 4, 9, 16, and 25. The full system is implemented on LLaVA-1.5-7B and compared against PruMerge, DivPrune, MQT-LLaVA, QueCC, TokenPacker, and Fourier-VLM on GQA, MMBench EN/CN, MME All, POPE F1, ScienceQA, VQA-Text, and MMVet, with all baselines sharing the same retraining/evaluation harness and matched token budgets.
Why This Matters
Impact on research. The paper argues that extreme compression is limited by two constraints — the retained subspace must concentrate task-relevant information, and the retained coordinates must induce a tractable alignment problem — and turns both into measurable functionals. This gives the field diagnostics that separate what is kept from how it is organized, rather than treating compression as a token-ranking score. It also positions itself against learned resamplers, which achieve accuracy but add attention computation, parameters, staged training, and alignment complexity; Braco's compressor cost is measured independently of downstream prompting.
Real-world applications:
- Mobile and on-device multimodal inference, where a fixed, lightweight visual interface is needed and compressor cost must be measured independently of the downstream prompt.
- Low-latency interactive agents, where the paper reports up to approximately 36% end-to-end speedup and single-image prefill latencies around 40–41 ms at 25/16/9/4 tokens.
- Long-context multimodal reasoning, where the visual interface grows and keeping the retained-token budget small converts extreme compression into larger end-to-end savings — for example, 40.57T to 3.45T FLOPs at 2880 input tokens.
- Cost-constrained deployment on smaller LLMs, where the paper reports Braco at 93.1/92.6 Acc. at 729/1024 tokens and 90.8 Acc. at 3645 tokens with Qwen2.5-3B, at 182× compression.
Industry relevance. The measured prefill FLOPs and latency (1.09–1.37T FLOPs, roughly 40–41 ms at 4–25 tokens versus 8.67T and 67.25 ms for the 576-token Vanilla model) map directly onto serving economics. The reported 16.6× lower compressor latency and 78.8× fewer compressor FLOPs relative to QueCC, plus lower module memory (8.731 MB versus 14.875 MB pre-projector), speak to deployment rather than only benchmark accuracy.
Future Directions
-
Extending the diagnostics to larger and more varied vision backbones and LLMs. The paper reports generalization across Vicuna-7B and Qwen2.5-3B with input lengths up to 3645 tokens; whether the budget-dependent coordinate rule (vanilla below K_b = 16, idct at or above) holds for other backbones is left to check.
-
Understanding the Qwen2.5-3B degradation of query-dependent compression. The authors suggest it may reflect sensitivity of query-dependent methods to weaker prompt understanding in smaller LLMs, which is a hypothesis rather than a settled explanation.
-
Pushing beyond 144× toward the smallest budgets. Braco stays within 0.2 Acc. of QueCC at 4 tokens, so the 4-token and beyond regime remains an open frontier where the residual-token budget and backbone split must be re-tuned.
-
Exploring alternative bases and embeddings. The ablations quantify costs for spatial/Haar substitutions (2.3–3.0 Acc. points), coordinate-organization substitutions (2.5–3.1 points), and embedding substitutions (0.8–0.9 points), which points to further design-space search over bases, embeddings, and coordinate organizations.
Target Audience
Researchers and engineers working on vision-language model efficiency, multimodal inference serving, and token compression. It is best suited to readers comfortable with transform coding, orthonormal bases, and optimization analysis, who want both a design method and a diagnostic framework for choosing visual interfaces under extreme token budgets. Practitioners deploying VLMs on latency- or cost-constrained hardware will find the compressor cost measurements and end-to-end prefill numbers directly applicable.
Authors’ abstract
Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under $23\times$--$64\times$ compression and remains competitive at $144\times$, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.