Skip to content
AI.info

Research

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Overview Research area: Computer vision and multimodal machine learning — specifically unified multimodal models (UMMs) that

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
arXiv
2609.38597
Published
2026-09-29
Authors
Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taixé, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu

AI summary

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

Overview

Research area: Computer vision and multimodal machine learning — specifically unified multimodal models (UMMs) that perform both visual understanding and visual generation.

Technical level: Advanced. The paper assumes familiarity with vision Transformers, VAEs, flow matching, mixture-of-experts style architectures, and large-scale multimodal training recipes.

Scope in one sentence: PixelUMM is an encoder-free model that represents images as 2D patches and videos as 3D tubelets, feeding raw pixels straight into a Mixture-of-Transformers backbone so that a single system handles image understanding, video understanding, text-to-image generation, and text-to-video generation.

What This Paper Is About

Unified multimodal models typically carry two separate visual interfaces: a vision Transformer (ViT) that produces semantic features for understanding, and a variational autoencoder (VAE) that produces reconstruction-oriented latents for generation. This roughly doubles the visual context compared with a single visual stream of similar token resolution, raising attention and memory costs, and it diverges from the single visual stream used in vision-language model pretraining, making unification at pretraining time impractical. Pixel-space modeling removes the latent encoding stage entirely, but extending it from images to video is non-trivial because video understanding and video generation conventionally adopt different temporal representations. PixelUMM's goal is to design a shared, encoder-free pixel-space interface that works for both.

Key Contributions

  1. An encoder-free unified model for image and video. PixelUMM contains no vision encoder, no VAE, and no discrete visual tokenizer. The only transformations between pixels and the backbone are one-layer linear projections plus deterministic patchify/unpatchify operations. Images are partitioned into non-overlapping 16×16 patches and videos into 4×16×16 tubelets, each mapped to the hidden size by exactly one linear layer.

  2. A unified visual interface spanning two temporal conventions. The paper introduces two video-understanding modes: dense_mode for inputs at least 4 FPS, which applies 3D patchify at a default 4 FPS sampling rate, and sparse_mode for inputs below 4 FPS, which samples at 1 FPS and encodes each frame independently with the image understanding projection. The same interface supports text-to-video generation and image-to-video conditioning.

  3. Joint text and pixel-space objectives in one backbone. A Mixture-of-Transformers (MoT) architecture pairs an understanding expert and a generation expert with symmetric structures, expert-specific normalization, projections, and FFNs, while all tokens interact through shared multimodal self-attention in every block. The model jointly learns autoregressive text prediction and pixel-space flow matching, with clean visual inputs routed to understanding and noisy visual inputs routed to generation. Explicit timestep embeddings and timestep-conditioned AdaLN are omitted.

  4. Controlled empirical studies of design choices. Eight experimental families (F1–F8) examine image patch size, video patch size, output decoder design, pixel-space versus VAE-space training dynamics, model size, compute scaling, and multimodal context conditioning.

Main Findings

  • Image patch size (F1): With the same sequence-length budget, 32×32 patches process four times as many images per step as 16×16 patches, yet the 16×16 run maintains lower image-generation training loss (T2I MSE) throughout late training. The conclusion is that stronger spatial compression makes image generation harder to learn. Qualitative comparisons use two matched prompts at 6K, 10K, and 15K steps with DPM-Solver at 50 steps, timestep shift 1, and CFG 3.5 with global renormalization.

  • Video patch size (F2): Four tubelet-to-token mappings were compared — p32/t4, p32/t2, p16/t4, and p32/t1. Less aggressive spatiotemporal compression generally yields lower T2V loss, with p32/t1 (F2-R04) attaining the lowest loss. The final model nonetheless adopts p16/t4 (F2-R03) to align with the spatiotemporal compression convention used by common video VAEs such as Wan2.2. Compression across tested configurations ranges from 1,024 pixels to one video token (p16/t4 or p32/t1) up to 4,096 pixels to one video token (p32/t4).

  • Patch artifacts and decoder heads (F3): Grid-aligned intensity changes appear in smooth, low-texture regions at high classifier-free guidance scales such as around 6 when using a linear pixel head. Three convolutional alternatives were tested at 96×176×320: a Wan-style upsample-conv decoder (F3-R02), a temporal-first PixelShuffle decoder (F3-R03), and a spatial-first PixelShuffle decoder (F3-R04). The linear head costs 133 GFLOPs with 12.6M parameters; the three convolutional heads cost 2,274, 1,279, and 747 GFLOPs (17.1×, 9.6×, and 5.6× the linear head) with 18.3M, 52.9M, and 15.1M parameters (1.5×, 4.2×, and 1.2× the linear head). The T-S and S-T PixelShuffle heads converge to losses near 0.02, while the Wan-style head remains above 0.05 over the observed interval. On boundary probes over 72 evaluation prompts (288 videos total), F3-R01 shows the strongest spatial and temporal boundary elevations (1.078 and 1.160); the Wan-style decoder is near the normalized baseline, while the PixelShuffle decoders retain smaller residual elevations. PixelUMM retains the linear heads for the released checkpoint.

  • Pixel-space versus VAE-space dynamics (F4): Over the first 10K steps, the logged VAE-space loss is about 4.3× the pixel-space loss, but the paper states this gap does not show that pixel-space training learns faster or produces better images because the losses are measured in different spaces. Pre-clip global gradient norms are nearly equal, while the pixel-space loss occasionally spikes.

  • Model size (F5): Under the same recipe and a global batch size of 256, the 8B model (F5-R02) reaches comparable losses in roughly one-third as many training steps as the 1.7B model (F5-R01): MSE 0.060 at about 14K versus 40K steps, and CE 0.50 at about 10K versus 31K steps — approximately 3× faster convergence in training steps, not wall-clock time.

  • Compute scaling (F6): With the same 1.7B model and matched data, visual interface, loss weights, and optimizer settings, moving from 8 to 128 GPUs benefits text convergence far more than generation: CE reaches 1.0 at about 2K versus 17.5K steps (roughly 9×), whereas MSE reaches 0.065 at about 17K versus 23K steps (roughly 1.3×). These are curve-based estimates of steps to a fixed loss, not wall-clock speedups.

  • Multimodal context conditioning (F7): PixelUMM extends encoder-free visual conditioning to image-to-video generation and video editing, routing clean image/video conditions through the understanding expert while noisy target tokens enter the generation expert and interact through shared attention. A multi-task fine-tuning stage on eight tasks for 15K steps starting from an intermediate Joint Stage 2 checkpoint produced mixed changes on image understanding benchmarks and improvement on all four video benchmarks. Image results included BLINK +2.37 and CV-Bench +2.27, while SEED-I decreased by 5.41 points. Video gains ranged from +0.89 to +3.49 points in dense_mode and +1.00 to +3.49 points in sparse_mode.

  • Overall performance: The abstract reports that PixelUMM achieves competitive performance across image and video understanding and generation tasks. The detailed main benchmark comparison in Section 3.2, and the contents of experimental family F8, are not included in the provided text.

Methodology in Plain English

PixelUMM starts from the observation that most unified models bolt two different visual encoders onto one language model. Instead, it feeds raw pixels directly into the Transformer and lets the same weights handle everything.

  • Input interface. An image is cut into 16×16 squares; a video is cut into 4-frame-by-16-by-16 blocks. Each block is flattened into a list of RGB values and passed through a single linear layer that projects it into the Transformer's hidden dimension. Separate linear layers exist for understanding and generation inputs, and clean inputs go through the understanding projections while corrupted targets go through the generation projections. At the output, a RMSNorm plus a linear layer maps hidden states back to RGB values for a patch or tubelet, which is then rearranged into pixels. These output projections are zero-initialized before multimodal training.

  • Two experts, one attention. Each Transformer block contains route-specific normalization, QKV/output projections, and FFNs for an understanding route and a generation route, but all tokens — text, clean visual, and noisy visual — participate in the same self-attention operation.

  • No timestep signal. Unlike models that feed an explicit timestep embedding into the denoising network, PixelUMM lets the network infer the corruption level from the noisy pixels themselves. The timestep is still used to build training targets and drive sampling.

  • Positional encoding. Following a three-axis Native RoPE design, half of each attention head is allocated to the temporal axis and one quarter each to height and width. Text tokens advance only along the temporal axis, while visual tokens also carry spatial grid coordinates.

  • Attention patterns. Conversations are serialized in ChatML. Attention is causal inside text but bidirectional inside each visual block, so an image is one bidirectional island and sparse video frames form temporally ordered islands where later frames can see earlier ones but not the reverse. For generation, the noisy target attends to the causal prompt and is bidirectional within the complete target block. Understanding masks are implemented with FlexAttention; generation uses two variable-length FlashAttention calls, one causal for the prompt and one non-causal for the target.

  • Objectives. Text supervision is cross-entropy applied only to assistant tokens, including the end-of-turn token, with square-root sequence-length normalization across responses. For pixels, a clean patch or tubelet is corrupted by linear interpolation with Gaussian noise, the head predicts the clean pixels, and the loss is computed in velocity space with a floor of 0.05 on the time denominator. The three losses are combined with stage-specific weights, and only loss terms relevant to each sample are active.

  • Staged training. Six stages progressively add tasks and resolution. Joint Stage 1 (150K steps) trains text-only, I2T, and T2I at 256×256 with loss weights 1:10:0. Gen Stage 1 (200K steps) trains T2I and T2V at 256×256 with weights 0:10:30. Und Stage 1 (100K steps) trains text and native-resolution image understanding. Und Stages 2 and 3 add video understanding at 224×224 for 20K steps and 448×448 for 15K steps. Joint Stage 2 (20K steps) trains everything together: T2I and T2V at 512×512, native-resolution image understanding, video understanding at 448×448, and text-only tasks, with weights 1:10:30. Learning rates are 1×10⁻⁴ for the first two stages and 2×10⁻⁵ afterwards, with constant schedules, AdamW, zero weight decay, gradient norm clipping at 0.2, and 1000 warmup steps. Video generation uses 4.0-second clips at 24 FPS (96 frames), and a time shift of sqrt(HW/(256×256)) applies only to

Authors’ abstract

Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.

Read the original paper