Research
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Overview Research area: Computer vision and multimodal foundation models — specifically the transfer of supervision from image-

- arXiv
- 2609.38079
- Published
- 2026-09-29
- Authors
- Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
AI summary
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?Overview
Research area: Computer vision and multimodal foundation models — specifically the transfer of supervision from image-to-image (I2I) generation tasks to image-to-text (I2T) visual understanding capabilities.
Technical level: Advanced. The paper assumes familiarity with Mixture-of-Transformers architectures, diffusion/flow-matching objectives, autoregressive cross-entropy training, gradient geometry, and benchmark-based evaluation.
Scope in one sentence: Through controlled paired tasks and a new 19-task / 25-capability taxonomy called OmniTaskonomy, the paper maps when and how visual generation training improves visual understanding, and links those gains to gradient alignment between the two objectives.
What This Paper Is About
Prior work has shown that visual understanding can improve visual generation, but the reverse direction is unclear, and some studies have concluded that generation provides substantially less benefit to understanding. That asymmetry is counterintuitive, because generation objectives provide dense pixel-level supervision over object appearance, spatial relationships, and geometry — precisely the signals that understanding tasks such as recognition, counting, spatial reasoning, and 3D perception rely on. The paper asks when and how generation supervision improves understanding, and builds a shared taxonomy so that transfer can be measured at the level of individual visual capabilities rather than whole benchmarks.
Key Contributions
-
A controlled paired-task study of the training recipe. The authors construct image-to-image and image-to-text versions of the same tasks (Jigsaw and Zoom-In, adapted from VisGym) that share the same input and the same underlying visual problem but differ in output modality, then compare six training recipes to isolate what enables transfer.
-
OmniTaskonomy, a unified taxonomy spanning generation and understanding. The taxonomy organizes 19 I2I generation tasks and 25 I2T understanding capabilities under the three Rs of computer vision — Recognition, Reconstruction, and Reorganization — with I2I tasks and understanding capabilities occupying separate leaves of the same hierarchy.
-
A full generation-to-understanding transfer map. The paper measures transfer from all 19 I2I tasks to the understanding capabilities across 9,444 evaluation examples, reporting which capabilities receive significant gains and from which sources, including non-obvious pairings.
-
An optimization-level explanation via gradient alignment. The authors measure alignment between I2I and I2T gradients, localize it to early pre-attention RMSNorm parameters of the understanding branch, and show it is positively correlated with downstream transfer.
Main Findings
-
Sequential I2I then I2T training is the recipe that scales. Training with an initial I2I stage that updates parameters shared with the I2T objective (I2I → I2T, and also I2I → Mixed) produces gains that grow consistently as I2I data increases. Directly mixing the two objectives from the outset, or freezing the shared parameters during I2I training, produces weaker or less stable gains.
-
I2I supervision substitutes partially for I2T supervision, and most when I2T data is scarce. With the I2I budget fixed at 100k examples, gains appear at all four I2T budgets tested (1k, 3k, 10k, 30k), with larger gains at smaller budgets. On Zoom-In, 100k I2I examples followed by only 1k I2T examples match training on 10k I2T examples alone. On Jigsaw, 100k I2I examples followed by only 3k I2T examples match training on 10k I2T examples alone.
-
Correspondence holds at the individual-example level. Averaged over three seeds trained on 30k I2I and 1k I2T examples, I2T accuracy is higher where the paired I2I prediction was correct. For Jigsaw, I2I accuracy is 70.9, I2T accuracy 81.3, P(I2T correct | I2I correct) is 88.2, and P(I2T correct | I2I incorrect) is 64.3. For Zoom-In, the corresponding values are 91.8, 90.6, 90.9, and 86.7.
-
Transfer is selective and concentrated. Of the 19 capabilities displayed in the transfer map (those with more than 100 evaluation examples), five improve significantly with at least one I2I source under the same recipe. Metric 3D relation and counting benefit from the broadest set of sources at 12 I2I tasks each; 2D ordering follows with eleven. Capabilities such as OCR, text recognition, and appearance understanding show no improvement from any tested source.
-
Shared visual operations predict several of the strongest gains. Localization and object pointing produce the largest improvements in counting (+2.5 and +2.0 percentage points). Z-depth, Euclidean depth, and surface normals improve metric 3D relation by 3.6, 3.8, and 3.4 percentage points. Jigsaw improves 2D ordering by 6.8 percentage points.
-
Useful transfer is not confined to closely matched capabilities. Inpainting improves counting (+1.5 pp) and 2D ordering (+7.2 pp) even though neither target explicitly requires reconstructing missing regions. 2.5D segmentation improves category recognition (+1.2 pp), and the authors also report that Z-depth prediction improves localization.
-
Gradient alignment localizes to early normalization layers. In the controlled Jigsaw and Zoom-In pairs, alignment between I2I and I2T gradients is strongest in the understanding branch's pre-attention RMSNorm parameters, and within those parameters it is strongest in the earlier transformer layers. For this analysis the authors sample 500 matched examples per task and compute gradients at the pretrained checkpoint.
-
Gradient alignment tracks transfer. Averaging over the 19 I2I sources, mean alignment and mean transfer across the seven understanding capabilities with more than 500 evaluation examples are strongly positively correlated (r = 0.795). Across all 19 × 7 = 133 individual source–target pairs, alignment and transfer are positively correlated (r = 0.529).
-
Human validation supports the taxonomy annotations. Four human annotators reviewed a stratified subset of 250 samples covering all 25 capabilities, each reviewing 50 shared examples plus 50 non-overlapping examples, yielding 400 human reviews. Among definite judgments, 97.4% of human labels match the VLM-assigned capability; Cohen's κ ranges from 0.956 to 0.989 across annotators (mean 0.972), and on the 50 shared examples Krippendorff's α = 0.943.
Methodology in Plain English
The authors work with "omni models," which learn visual generation and visual understanding inside a single model, making transfer between the two directly observable. They use BAGEL (specifically BAGEL-7B-MoT) as their base, a Mixture-of-Transformers architecture whose generation and understanding branches have separate projections, normalization layers, and MLPs but interact through shared attention; the VAE encoder and decoder stay frozen throughout all training.
To isolate what causes transfer, they first build matched task pairs. In Jigsaw, patches of an image are shuffled and must be put back in order; the I2I version outputs the correctly ordered image, the I2T version outputs the same ordering as a permutation in text. Zoom-In is analogous but with views at different zoom levels. Because input and underlying problem are held fixed, differences in downstream I2T performance can be attributed to the generation objective. They compare six recipes — I2T-only, I2I → I2T, Mixed → I2T, Frozen I2I → Mixed, Mixed, and I2I → Mixed — with the I2T budget fixed at 1k examples and the I2I amount varied across pool sizes of 1,876, 9,380, 30,016, and 100,000 examples.
To move beyond paired tasks, they build OmniTaskonomy. They sample questions from seven vision-language benchmarks (BLINK, MMStar, MMT-Bench, CV-Bench, RealWorldQA, MMVP, and VStarBench), use gemini-3-flash-preview to describe the visual attributes each question requires, consolidate those descriptions into a fixed attribute schema, and annotate the full collection consistently. A VLM then groups annotated samples into local task trees that are merged into a global tree, with manual review to split, merge, and sharpen leaf boundaries. The resulting fine-grained capabilities are placed under Recognition, Reconstruction, or Reorganization. I2I tasks are placed in the same hierarchy according to the capability their objective directly supervises. Finally, three VLM judges independently assign each eligible benchmark example to one capability, and an example is kept only when at least two judges agree, yielding 9,444 samples. Human review checks the automated assignments.
For the transfer experiments, they start from the same BAGEL checkpoint, train on approximately 50k I2I examples and then 50k LLaVA-Instruct examples, and compare against a baseline trained only on the I2T data. Transfer from source s to capability t is the accuracy difference on capability t, in percentage points, and significance is assessed with exact paired permutation tests at p < 0.05. Generation outputs are scored by a VLM judge, gemini-3.7-flash, which must confirm all positions match, each source piece appears exactly once, task-fulfillment and content-fidelity scores are at least 3 on a 0–4 scale, and no decisive failure is reported.
To explain the transfer pattern, they measure whether the two objectives push the same parameters in the same direction. For each paired example, they compute the gradient of each objective's loss with respect to a parameter block at the pretrained checkpoint, normalize it, project both normalized gradients onto a shared uncentered PCA basis (retaining the fewest components whose eigenvalues sum to at least 99% of the total), take the cosine similarity in that subspace, and scale it by the square root of the estimated effective dimension. Alignment is then compared with downstream transfer.
Why This Matters
Impact on research. The paper challenges the view that generation is largely a one-way beneficiary of understanding. It reframes the question from "does generation help understanding" to "which generation tasks help which understanding capabilities, and under what curriculum," and it offers gradient alignment as an optimization-level signal that could be used to predict transfer before running expensive downstream training. The unified taxonomy also provides a shared vocabulary for comparing generation and understanding tasks that existing benchmarks, organized by output modality or task formulation, do not provide.
Real-world applications (implied by the findings):
- Training multimodal assistants where labeled understanding data is scarce and abundant unlabeled or synthetic images can be turned into I2I supervision — the paper shows I2I supervision can partially substitute for I2T data.
- Designing data-collection priorities for models that must count objects or reason about 3D relations, since counting benefits most from localization and object pointing, and metric 3D relation benefits most from depth and surface-normal prediction.
- Selecting auxiliary generation objectives when fine-tuning a model for a specific capability, using the transfer map as a lookup rather than relying on intuition about which tasks are "close."
- Building training pipelines for unified models in which the generation pathway supplies dense supervision that shapes the shared representation used by the understanding pathway.
Industry relevance. Teams building unified or omni models must decide how to allocate training compute and data between generation and understanding objectives. This paper provides evidence about curriculum order (an initial I2I stage that updates shared parameters), about which generation objectives are worth the compute for a given target capability, and about a cheap gradient-based diagnostic that could be run at the pretrained checkpoint to screen candidate task pairings before committing to full training runs.
Future Directions
-
Test whether gradient alignment predicts transfer for task pairs and models not studied here. The correlations reported (r = 0.795 across seven capabilities, r = 0.529 across 133 pairs) are associations, not demonstrations that alignment causes transfer; causal tests would require intervening on alignment during training.
-
Explain the capabilities that resist transfer. OCR, text recognition, and appearance understanding showed no improvement from any of the 19 tested I2I sources, and the transfer map is highly non-uniform overall. Understanding why these capabilities are immune could reveal what generation supervision cannot supply.
-
Extend the taxonomy beyond the current scope. OmniTaskonomy covers 19 I2I tasks and 25 understanding capabilities, derived from seven vision-language benchmarks. Whether the same structure holds for video, 3D, or interactive settings, and how the taxonomy should evolve as new tasks appear, is not addressed.
-
Turn the transfer map into an automated curriculum. The paper argues that generation objectives should be chosen for the visual capabilities they support; a natural next step is to select generation task mixtures automatically for a target understanding capability, potentially guided by measured gradient alignment rather than by manual inspection of the map.
-
Determine how far the recipe findings generalize. The controlled scaling results and the transfer experiments are conducted on BAGEL under a specific I2I → I2T recipe with fixed budgets (approximately 50k I2I and 50k LLaVA-Instruct examples in the taxonomy experiments); the paper does not report whether the same recipe ordering, or the same alignment-transfer relationship, holds for other omni architectures or at larger scale.
Target Audience
Researchers and engineers working on unified multimodal or omni models, multimodal pretraining, and transfer learning will benefit most, particularly those deciding how to combine generation and understanding objectives in a training pipeline. The paper is also relevant to readers interested in visual task taxonomies and evaluation design, and to those studying optimization-level diagnostics such as gradient conflict and task affinity, since it extends that line of work from multi-task supervised learning into the generation-versus-understanding setting. Readers seeking an introductory treatment of these areas will find the paper assumes substantial background.
Authors’ abstract
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.