Skip to content
AI.info

Research

QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

Overview Research area: Computer vision, specifically visual tokenization and autoregressive image generation (plus image compression and controllable synthesis). Technical level: Advanced. The paper

QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
arXiv
2610.10497
Published
2026-10-07
Authors
Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Divyansh Srivastava, Bingnan Li, Zhuowen Tu

AI summary

Overview

Research area: Computer vision, specifically visual tokenization and autoregressive image generation (plus image compression and controllable synthesis).

Technical level: Advanced. The paper assumes familiarity with VQ-VAE / VQ-GAN tokenizers, vector quantization, Transformer attention masking, quadtree data structures, and standard generative metrics (FID, PSNR, PSNR, Inception Score, precision/recall).

Scope in one sentence: QuadTok proposes a quadtree-structured visual tokenizer that spends tokens where an image is visually complex while keeping explicit token-to-region correspondence, and shows that the resulting coarse-to-fine token sequences can be generated autoregressively — including zero-shot spatial layout control.

What This Paper Is About

Standard visual tokenizers compress an image into a rigid 2D grid with a fixed token budget, spreading tokens uniformly even though natural images have highly non-uniform information density — complex textures need more capacity, smooth regions need less. Recent 1D tokenizers allow variable-length sequences but lose explicit spatial grounding, so tokens no longer map cleanly onto image regions. QuadTok's goal is to get both: adaptive, variable-length token allocation driven by local visual complexity, while preserving an explicit correspondence between each token and the image patch it represents, so the resulting sequence can drive autoregressive image generation.

Key Contributions

  1. A quadtree-based visual tokenizer (QuadTok) that replaces fixed-grid representations with spatially adaptive token allocation. Each node of the quadtree is a token tagged with a level l and spatial index p, and the tree is serialized by breadth-first search into a 1D sequence for Transformer processing, preserving explicit spatial correspondence.

  2. Region-Wise Complexity Guidance, a data-driven strategy that selects a content-adaptive tree per image based on reconstruction benefit. Two complementary probing trees (one expanding half the coarse regions, one expanding the other half) let each region be observed once expanded and once collapsed, and the averaged LPIPS loss difference b_r decides whether a region with b_r > τ is expanded into four children.

  3. A Kinship Causal Mask, a topology-aware attention mask used inside the tokenizer that enforces a strict coarse-to-fine dependency among quadtree tokens: parents attend to each other bidirectionally, children attend to parents (not vice versa), and children attend bidirectionally only to siblings sharing the same parent.

  4. Quadtree-conditioned autoregressive generation, where a decoder-only Transformer predicts visual codes in breadth-first order under a topology supplied before generation — enabling competitive class-conditional ImageNet generation and novel zero-shot spatially controlled synthesis.

Main Findings

  • Adaptive tokenization saves tokens at comparable fidelity. On ImageNet with an average of 230 tokens, the tokenizer reaches rFID 1.46 and PSNR 20.37, saving roughly 10% relative to a fixed 256-token grid. Transferred zero-shot to COCO with 232 tokens on average, it reaches rFID 7.88 and PSNR 19.30, saving roughly 9%.

  • The tradeoff versus fixed grids. Compared with LlamaGen at downsampling factor 16 (256 tokens, rFID 2.19, PSNR 20.79 on ImageNet), content-adaptive allocation achieves lower rFID with fewer tokens, at the cost of lower PSNR; this tradeoff persists under zero-shot transfer to COCO.

  • Competitive generation from a 947M model. QuadTok-XXL (947M parameters) reaches gFID 2.08, IS 273.35, precision 0.82, recall 0.59 on ImageNet 256×256, compared with 2.15 gFID for the 1.4B RandAR-XXL. Smaller variants: QuadTok-B (111M) gFID 4.41, QuadTok-L (343M) gFID 2.93, QuadTok-XL (697M) gFID 2.23.

  • Consistent scaling behavior is reported across the 111M, 343M, 697M and 947M two-level generators, plus a 344M three-level generator (gFID 2.77).

  • The Kinship Causal Mask is essential for spatial structure, not for rFID alone. Removing it slightly lowers rFID (1.46 → 1.41) but drops PSNR by over a point (20.37 → 19.16), and with the same 697M generator and CFG it raises gFID from 2.23 to 3.17. A single perturbed fine-level token stays spatially localized with the mask and spreads across the image without it.

  • Complexity guidance beats random and full expansion. A fully expanded tree needs 320 tokens (rFID 1.50, PSNR 20.39); a random tree of similar size needs 253 tokens (rFID 1.601, PSNR 19.40); guided QuadTok matches full-tree fidelity (rFID 1.46, PSNR 20.37) with 230 tokens, i.e. 28% fewer tokens than the full tree. Default guidance uses LPIPS with τ = 0.05.

  • Deeper trees improve reconstruction. A three-level configuration (adding a 32×32 token grid to the default 8×8 → 16×16 hierarchy) reaches rFID 0.770 and PSNR 23.31 at 989 tokens on ImageNet, and rFID 3.92 / PSNR 23.83 at 1012 tokens on COCO. Depth and token budget increase together in this experiment.

  • Zero-shot spatial layout control works with a frozen model. Using 192 tokens and 1,000 images per condition, prescribed topologies beat random topologies at the same token budget on all four target half-planes (left, right, top, bottom) across Grad-CAM, Grounding DINO box, and SAM 2 mask center metrics. Largest gains: left-plane Grad-CAM 29.6 → 58.3 (+28.7) and right-plane Grad-CAM 70.4 → 87.9 (+17.5). FID-50cls (against 2,500 class-matched real images) moved between −0.15 and +0.12 points relative to the matched random baseline.

  • Expansion is not monotonically beneficial. Sweeps over quadtree expansion probability (two-level p, and three-level p₂ with p₁ = 0.95) show gFID improving initially with expansion but reversing at higher expansion, so best evaluated quality occurs before maximal expansion.

  • Latency reduction is claimed but not quantified in the main text. QuadTok is reported to achieve lower generation latency than LlamaGen at 256×256 for matched model scales, with the timing protocol deferred to Appendix E.

Methodology in Plain English

  1. Represent the image as a tree. Split the image into a coarse grid of regions; each region is a token. If a region is refined, it becomes four child tokens covering its quadrants. Two levels are the default (an 8×8 grid and a 16×16 grid); a three-level version is also tested.

  2. Turn the tree into a sequence. Flatten it breadth-first, so all coarse tokens come first and their finer children follow. This ordering naturally matches coarse-to-fine generation.

  3. Encode. A ViT extracts dense patch features. Each tree node also gets a learnable instruction token encoding its level and spatial position. An Aggregator Transformer lets each instruction token read from the globally contextualized ViT features, producing one embedding per node.

  4. Quantize. Node embeddings are mapped to the nearest entry in a shared codebook of 16,384 entries, giving discrete visual tokens.

  5. Train with random trees. Rather than always using a fully expanded tree, each training iteration randomly expands parent nodes with probability γ = 0.5, keeping all level-1 tokens. This prevents the model from depending on one fixed layout.

  6. Constrain attention. The Kinship Causal Mask restricts interactions among tree tokens during encoding and decoding to parents, ancestors, and siblings, so refinement stays inside its own spatial region instead of leaking across the image.

  7. Choose which regions to refine. At tokenization time, reconstruction benefit is estimated via two complementary probing trees and a spatial LPIPS loss map; regions whose benefit exceeds τ are expanded.

  8. Decode hierarchically. Decoded node features are unpatchified with level-specific operators, placed on a canvas, and combined level-by-level through learned upsampling plus residual addition, starting from a coarse global map (H₀ = 0).

  9. Generate. A LLaMA-style decoder-only Transformer with a standard causal mask (not the Kinship mask) predicts visual codes in BFS order, conditioned on the class label and a topology fixed before sampling — so a user-chosen or algorithmically constructed tree can steer the output.

Why This Matters

Impact on research. QuadTok sits between two existing paradigms — rigid 2D grid tokenizers that preserve spatial grounding but waste tokens, and 1D tokenizers that allow variable length but are spatially agnostic. If adaptive token counts with explicit spatial binding hold up, it changes how tokenizers are designed for autoregressive image models: instead of choosing between grids and sequences, the representation itself becomes spatially adaptive. The result that the Kinship mask barely matters for reconstruction rFID but substantially matters for generation gFID (2.23 vs 3.17) is a pointed finding about what spatial structure buys an autoregressive model.

Real-world applications (as suggested or enabled by the paper's claims):

  • Content-adaptive image compression, where bits are spent on textured detail and smooth regions use coarser codes.
  • Controllable image synthesis, where a user or an LLM specifies a subject region that is converted into a quadtree, and the frozen generator places the subject there without layout-specific training.
  • Layout-aware content creation pipelines, where a spatial layout is translated into a tree and used as a conditioning interface before sampling.
  • Efficient high-resolution generation, since token budget scales with image complexity rather than pixel count — the paper reports 512×512 reconstruction and generation in Appendix D using 16×16 and 32×32 grids, with fewer tokens than 1,024-token grid baselines.

Industry relevance. Autoregressive image generators are compute- and sequence-length-bound; a tokenizer that averages 230 tokens instead of 256 (and reports lower latency at matched scales) directly affects serving cost. The layout interface is also a natural hook for multimodal and text-to-image pipelines, which the authors explicitly leave to future work.

Future Directions

  • Text-to-image and multimodal generation. The paper states its scope is class-conditional generation and that QuadTok is compatible with multimodal pipelines and text-to-image generation, which is left for future work.
  • Better region-selection criteria. The default uses LPIPS-based guidance with τ = 0.05; the paper defers a full probing-metric and threshold ablation to Appendix G, and the expansion sweeps show that more expansion is not always better, so optimal tree selection remains open.
  • Extending depth and resolution. Three-level trees and 512×512 evaluation are demonstrated, but the paper notes that depth and token budget increase together in that experiment, leaving the question of how to trade depth against budget.
  • Learning topologies rather than supplying them. In class-conditional evaluation the topology is sampled independently of image content, and for layout control it is user-specified; predicting the tree itself (rather than receiving it) is an unaddressed direction.

Target Audience

Researchers and practitioners working on visual tokenizers, discrete latent representations, and autoregressive or next-token image generation; engineers building image generation or compression systems who care about token budget and spatial controllability; and readers interested in controllable or layout-conditioned synthesis. Readers unfamiliar with vector quantization and Transformer attention masking will find the method sections demanding, while the reconstruction and generation result tables are directly comparable against the listed baselines.

Authors’ abstract

We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet $256 \times 256$ benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: https://github.com/myc634/QuadTok.

Read the original paper