Research
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Overview Research area: Computer Vision — specifically visual tokenizers/autoencoders for latent image generation and text-to-image synthesis. Technical level: Advanced. The paper assumes familiarity

- arXiv
- 2609.37775
- Published
- 2026-09-29
- Authors
- Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou
AI summary
Overview
- Research area: Computer Vision — specifically visual tokenizers/autoencoders for latent image generation and text-to-image synthesis.
- Technical level: Advanced. The paper assumes familiarity with Vision Transformers, latent diffusion, Fréchet distances, and reconstruction metrics such as PSNR and LPIPS.
- Scope: The paper introduces HiRAE, a tokenizer that fuses all 24 layers of a frozen DINOv3-L encoder into a single latent, and evaluates it on ImageNet-256 reconstruction and generation plus text-to-image alignment benchmarks.
What This Paper Is About
Frozen pretrained vision encoders produce representations that are good for generation but can drop the fine visual detail needed to reconstruct an image faithfully. Intermediate layers of those same encoders hold much of that missing detail, but fusing layers together tends to distort the latent space the generator must model, and prior methods either hand-pick which layers to use or train fusion and decoding in separate stages. HiRAE learns to combine the full encoder hierarchy in one joint training stage, using depth-dependent limits on how much each layer group is allowed to change the deepest-layer representation.
Key Contributions
- Hierarchical representation autoencoding with depth-dependent residual budgets. HiRAE groups encoder layers into shallow, middle, and deep sets, and each group learns a residual correction to the deepest feature under a separate norm cap that is tighter for shallower groups.
- Higher reconstruction fidelity plus improved text-to-image alignment. HiRAE-24 reduces reconstruction FID by roughly 30% relative to RAEv2 (0.299 to 0.209) while improving GenEval, DPG-Bench, and GenAI-Bench before and after supervised fine-tuning, including a 2.84-point GenEval gain after fine-tuning.
- Analysis of HiRAE's latent structure and decoding sensitivity. The paper reports that fusion enriches spatial detail while largely preserving class neighborhoods, and that the learned tokenizer shows lower decoding sensitivity to the tested latent perturbations.
- Removal of manual layer selection and of a separate fusion-only adaptation phase. HiRAE-24 learns contributions from all 24 layers, while HiRAE-7 shows the framework can also extend an established seven-layer subset.
Main Findings
- Reconstruction on ImageNet-256: HiRAE-24 reduces rFID from the RAEv2 result of 0.299 to 0.209, described as a reduction of approximately 30%, using the same frozen DINOv3-L encoder and a 16×16×1024 latent.
- Matched 5,000-image subset: PSNR rises from 22.667 to 26.377 dB and LPIPS falls from 0.074 to 0.043, covering both pixel accuracy and perceptual similarity.
- Guided ImageNet generation: After 80 epochs of generator training, guided gFID drops from 1.060 (RAEv2) to 1.038, Inception Score rises from 255.300 to 257.823, and FD_r^6 improves from 2.170 to 1.856.
- Unguided generation: HiRAE-24 reaches gFID 2.129, improving on RAEv2 K=23's 3.010 but remaining above RAEv2's 1.650; IS is 210.339 (RAEv2 K=23: 206.000) and FD_r^6 is 4.660 (RAEv2: 3.950).
- Text-to-image after pretraining: HiRAE-24 improves GenEval by 4.52 points and DPG-Bench by 1.51 points over RAEv2. Reported scores are GenEval 60.94, DPG-Bench 82.73, GenAI-Bench 68.08 for HiRAE-24, versus 56.42, 81.22, and 67.02 for RAEv2.
- Text-to-image after SFT: Gains of 2.84, 1.45, and 0.97 points on GenEval, DPG-Bench, and GenAI-Bench. GenEval rises from 84.86 (RAEv2) to 87.70 (HiRAE-24), with DPG-Bench 86.35 versus 84.90 and GenAI-Bench 72.66 versus 71.69.
- HiRAE-7 also helps: On the seven selected layers, HiRAE-7 reduces rFID from 0.299 to 0.217 and guided gFID from 1.060 to 1.038, and exceeds RAEv2 on every available text-to-image benchmark.
- Reconstruction fidelity alone does not predict generation quality: DRoRAE-style fusion over all 24 layers reaches a lower rFID (0.065) than HiRAE-24 (0.209) but a worse guided gFID (1.551 versus 1.038) and FD_r^6 (3.375 versus 1.856).
- Ablation on expert inputs and regularization: Replacing 24 raw-layer experts with seven depth-mode experts raises rFID from 0.209 to 0.230 and epoch-80 guided gFID from 1.038 to 1.067. Removing residual regularization reaches rFID 0.023 but gFID 7.905 and FD_r^6 14.722 at epoch 20.
- Depth-group count: Three groups achieve the lowest guided gFID in that comparison. Reported values are rFID 3.293/guided gFID 29.385 for two groups, 3.363/28.354 for three, and 3.457/31.321 for four.
- Depth-group interventions: Removing any group's residual increases reconstruction LPIPS on 5,000 matched images. Reported LPIPS changes (×10^3) for group removal are 5.650 for shallow (0–7), 54.111 for middle (8–15), and 147.750 for deep (16–23).
- Latent organization largely preserved: In the original feature space, the same-class fraction among ten nearest neighbors over 5,000 matched images is 74.454% for RAEv2 and 79.536% for HiRAE-24. Within HiRAE-24 this fraction changes from 79.634% for the deep anchor to 79.536% after fusion, and mean spatial CKA remains 0.985.
- Spatial variation spreads: Across 100 fixed images, spatial effective rank increases from 130.360 for the deep anchor LN(H_23) to 154.893 for the fused representation, with an increase in every image.
- Lower decoding sensitivity: At a perturbation norm equal to 10% of the standardized latent norm, HiRAE-24 produces 21%–22% of RAEv2's output LPIPS change across the three tested direction types.
Methodology in Plain English
HiRAE starts from a frozen DINOv3-L vision transformer and keeps its 24 intermediate layer outputs. Rather than hand-picking which layers feed the tokenizer, it learns a small MLP "expert" for each layer, plus a learned router that reads the deepest feature and assigns each spatial location a signed weight per layer.
The design choice that distinguishes HiRAE is how those weighted contributions are allowed to change the latent. The deepest layer's feature (H_23) is kept as an anchor. Layers are split into three depth groups — shallow (0–7), middle (8–15), and deep (16–23) — and each group's weighted sum becomes a residual correction to that anchor. Each group's correction is capped in Frobenius norm relative to the anchor, with caps of 0.025, 0.075, and 0.150 respectively, summing to an overall bound of 0.250. Shallow groups also get stronger residual dropout (0.50, 0.25, 0.10). This keeps shallow detail from dominating the latent. The final latent is the layer-normalized sum of the anchor and the three controlled corrections, preserving the original token count and channel dimension.
Training happens in two stages. Stage 1 jointly trains the fusion module and the decoder with pixel reconstruction, perceptual, and adversarial losses while the backbone stays frozen — no separate fusion-only adaptation phase. Stage 2 freezes the tokenizer and trains a DiT generator on its latents, following RAEv2's prediction and internal-guidance framework. For text-to-image, the team pretrains for 100K optimizer updates and then applies supervised fine-tuning for 2,850 updates on 8× H800 80GB GPUs.
Why This Matters
- Research impact: The paper shows that reconstruction fidelity and generation quality can diverge — a lower rFID (0.065 versus 0.209) did not translate into better guided generation — and offers a controllable mechanism (depth-dependent residual budgets) for balancing the two rather than relying on manual layer-subset search or staged training.
- Real-world applications:
- High-fidelity image compression and tokenization pipelines that need to preserve fine structures such as text strokes and local color.
- Text-to-image systems that must follow compositional prompts accurately, as measured by GenEval, DPG-Bench, and GenAI-Bench.
- Class-conditional image generation where training compute is limited, since HiRAE-24 reports competitive guided results at 80 generator epochs.
- Editing or restoration workflows where the tokenizer's decoding stability under latent perturbations matters.
- Industry relevance: The approach removes configuration effort — no layer-subset selection, no separate fusion-only adaptation stage — and preserves the latent shape (16×16×1024) of the base model, so it can be dropped into an existing representation-autoencoder training recipe without changing downstream generator architecture.
Future Directions
- Reconciling the DRoRAE trade-off. DRoRAE-style full-depth fusion achieved rFID 0.065 versus HiRAE-24's 0.209 but a worse guided gFID (1.551 versus 1.038); the paper does not report a configuration that attains both.
- Improving unguided generation. HiRAE-24's unguided gFID of 2.129 improves on RAEv2 K=23's 3.010 but remains above RAEv2's 1.650, an open gap the paper reports without resolving.
- Generalizing beyond the tested setup. All results use DINOv3-L/16 at 256×256 with three depth groups; the paper does not report whether the budget schedule transfers to other backbones, resolutions, or token counts.
- Refining the middle depth group. The middle group's removal difference is described as small, with a paired 95% interval that includes zero, leaving its role less clearly established than the shallow and deep groups.
Target Audience
Researchers and engineers working on visual tokenizers, latent diffusion, and representation autoencoding, particularly those building text-to-image or class-conditional generation systems that sit on top of frozen pretrained vision encoders. The paper is also relevant to practitioners who want to avoid manually tuning encoder layer selection, and to anyone studying the relationship between latent-space structure and generative modelability.
Authors’ abstract
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.