Skip to content
AI.info

Research

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Overview Research area: Unified multimodal generative modeling — specifically, a fully continuous approach to jointly

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
arXiv
2609.40362
Published
2026-09-30
Authors
Hongyuan Tao, Xinggang Wang, Lianghui Zhu, Yongkang Li, Yunchao Wei, Bin Feng, Shaoyu Chen, Qian Zhang, Chang Huang, Kai Yu

AI summary

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

Overview

Research area: Unified multimodal generative modeling — specifically, a fully continuous approach to jointly pretraining language and vision models in their embedding spaces (computer vision / multimodal foundation models, arXiv 2609.40362v1 [cs.CV], 30 Sep 2026, from Huazhong University of Science and Technology, Beijing Jiaotong University, and Horizon Robotics).

Technical level: Advanced. The paper assumes familiarity with Flow Matching, diffusion sampling, classifier-free guidance (CFG), autoregressive versus discrete-token modeling, and multimodal embedding encoders.

Scope in one sentence: The paper proposes Multimodal Flow, a framework that represents text blocks and images as ordered continuous "hyperchunks" and generates both with a single shared chunk-causal Flow Matching backbone, instantiated as MF-1 at 0.6B, 1.2B, and 1.6B scales.

What This Paper Is About

Most unified multimodal models either quantize images into discrete visual tokens so language and vision share one categorical sequence, or they keep discrete language prediction and pair it with continuous image diffusion/flow — meaning the two modalities are still governed by different objectives and sampling procedures. Multimodal Flow asks whether language and vision can instead share one fully continuous generative process in embedding space, while each modality keeps its own representation structure. The goal is to preserve visual fidelity without a tokenizer bottleneck, and to unify understanding and generation under a single objective and sampling mechanism.

Key Contributions

  1. A fully continuous multimodal framework. Multimodal Flow models language and vision in their respective continuous embedding spaces under one Flow Matching objective, avoiding visual quantization and avoiding modality-dependent objectives.

  2. Ordered hyperchunks and a chunk-causal backbone. Text blocks (8 tokens per block in the main models) and images (a 16×16 visual grid) become continuous chunks that retain textual token order and visual spatial structure. A chunk-causal flow backbone conditions each target chunk on preceding chunks, predicting multiple target chunks in parallel during training and generating chunks sequentially at inference.

  3. Joint attention with modality-specific computation. Joint attention provides cross-modal interaction while modality-specific feed-forward networks adapt computation to each representation space; the paper reports that sharing FFNs degrades performance, especially on image-conditioned text generation.

  4. MF-1 instantiation with scaling, controlled comparisons, and transfer evidence. MF-1 is pretrained on mixed multimodal data at 0.6B, 1.2B, and 1.6B scales, evaluated against generation-only, understanding-only, and unified models, compared against matched-budget Transfusion-style and Chameleon-style architectures, and tested for downstream transfer from mixed pretraining.

Main Findings

  • Generation performance at 150B tokens: With 150B pretraining tokens, the 1.6B MF-1 scores 0.821 on GenEval and 83.44 on DPG-Bench, which the abstract reports as an average of 82.8 across GenEval and DPG-Bench.

  • Understanding performance at 150B tokens: MF-1 reaches POPE 86.1, MMBench 67.2, SEEDB 62.4, VQAv2 72.6, GQA 58.3, and OK-VQA 39.1, which the abstract reports as an average of 75.3 across VQAv2, MMBench, and POPE.

  • Strongest compositional generation among unified models on GenEval: MF-1 records Single 0.99, Two 0.92, Counting 0.73, Colors 0.90, Position 0.81, Color Attribution 0.58, Overall 0.82, compared with overall scores of 0.67 for Transfusion (7B), 0.65 for D-DiT (2B), 0.63 for Hunyuan-DiT (1.5B) and JanusFlow (1.3B), 0.62 for SD3 Medium (2.0B), 0.61 for Muddit (1B), 0.55 for SDXL (2.6B), 0.54 for Emu3-Gen (8B), 0.53 for Show-o (1.3B), and 0.47 for LWM (7B).

  • Best DPG-Bench overall among the compared unified models: MF-1 records 89.38 Global, 89.03 Entity, 90.82 Attribute, 89.73 Relation, 89.47 Other, 83.44 Overall, compared with 80.09 for JanusFlow (1.3B), 80.60 for Emu3-Gen (8B), 79.68 for Janus (1.3B), 78.87 for Hunyuan-DiT (1.5B), 74.65 for SDXL (2.6B), and 67.27 for Show-o (1.3B); SD3 Medium (2.0B) reaches 84.08.

  • Competitive despite far fewer tokens and no pretrained LLM initialization: Trained from scratch on 150B tokens, MF-1 surpasses from-scratch unified models such as Muddit (1.0B, ~400B PT tokens: POPE 52.5, MMB 28.4, SEEDB 13.0, VQAv2 68.2, GQA 57.5, OK-VQA 21.1) and D-DiT (2.0B, ~4.3T PT tokens: POPE 79.2, VQAv2 59.5, GQA 55.1, OK-VQA 28.5), while remaining competitive with similarly sized models initialized from pretrained LLMs, such as Janus (1.3B, ~0.9T: POPE 87.0, MMB 69.4, SEEDB 63.7, VQAv2 77.3, GQA 59.1) and JanusFlow (1.3B, ~2.0T: POPE 88.0, MMB 74.9, SEEDB 70.5, VQAv2 79.8, GQA 60.3). Janus-Pro (1.5B, ~1.1T) reports POPE 86.2, MMB 75.5, SEEDB 68.3, GQA 59.3; Show-o2 (1.5B, ~18.2T) reports POPE 83.0, MMB 67.4, SEEDB 65.6, GQA 60.0.

  • Controlled architecture comparison at matched budget: Under matched data, optimization schedules, and trainable-parameter budgets (1.6B, 50B pretraining, 5B finetuning), Multimodal Flow (fully continuous) achieves GenEval 0.7134, GQA 55.60, VQAv2 69.03, MMBench 46.74, SEEDB 51.85, versus Transfusion-Style (discrete–continuous hybrid) 0.6693, 52.81, 68.49, 33.68, 31.31 and Chameleon-Style (fully discrete) 0.3744, 45.83, 56.37, 38.40, 42.40.

  • Mixed pretraining transfers: At matched total training-token budgets for the 1.6B architecture, mixed pretraining raises SeedBench from 31.6 (random initialization) to 62.4 and MMBench from 36.0 to 67.2; GenEval rises from 0.527 to 0.821, DPG-Bench from 57.81 to 83.44, and OK-VQA from 22.9 to 39.1. The random counterpart is trained directly on downstream data for 155B tokens, matching the pretrained model's combined pretraining (150B) and finetuning (5B) budget.

  • Scaling behavior: The 0.6B, 1.2B, and 1.6B variants all show the Flow Matching objective and language PPL decreasing while GenEval and CLIPScore increase; larger capacity yields clearer gains in language modeling and image captioning, with the 1.6B model strongest later in training.

  • Semantic visual representations win: DINOv2 achieves the highest image-generation scores, while SigLIP2 gives a better generation–captioning balance with higher CIDEr and CLIPScore; reconstruction-oriented VAE latents (FLUX.2 VAE, SD-VAE) underperform semantic embeddings on both generation and captioning, and raw image patches also perform weakly.

  • Modality-specific FFNs matter: After 50B pretraining tokens, sharing FFNs degrades results (Shared/Shared: Text PPL 28.35, GenEval 0.237, DPG 69.70, CIDEr 42.02, CLIPScore 0.788) relative to shared attention with specific FFNs (Text PPL 27.43, GenEval 0.226, DPG 71.04, CIDEr 47.91, CLIPScore 0.820), with the largest drop on image-conditioned text generation.

  • The same CFG mechanism works for both directions: Image captioning performs best at a guidance scale of 3, while text-to-image generation peaks at a scale of 5 and declines as the scale increases further.

Methodology in Plain English

The method starts by turning each modality into continuous vectors rather than discrete symbols. A frozen T5-small encoder converts text into 512-dimensional latents, processed in blocks of 8 tokens; a frozen SigLIP2-so400m encoder converts a 224×224 image into 256 patch embeddings of dimension 1152 that form a single spatially structured chunk. Both sets of vectors are normalized (text with dimension-shared constants of mean 0 and standard deviation 0.2; vision with per-token, per-channel statistics estimated from 50,176 GPIC training images) and projected into a common hidden space.

These normalized embeddings are called hyperchunks, and they are arranged into ordered sequences — for example, question chunks followed by answer chunks, or prompt chunks followed by an image chunk. A single chunk-causal transformer backbone processes the sequence: when predicting a chunk, all earlier chunks are visible and future chunks are hidden, but positions within the current chunk are modeled jointly. Information flows between modalities through joint attention, while modality-specific feed-forward networks handle each modality's statistics, and multimodal rotary position embeddings encode both chunk order and intra-chunk structure.

Training uses Flow Matching. Each target chunk is mixed with noise along a linear path, and the backbone predicts the clean endpoint, which is converted into a velocity and trained with a velocity mean-squared error. Because of the chunk-causal mask, many target chunks can be predicted in one forward pass, each with its own independently sampled timestep (shifted logit-normal with α=8 for image targets and α=6 for text targets). Independent examples are packed into physical sequences of at most 32,768 model positions, with sequence identifiers preventing attention across examples.

Pretraining mixes four task types: text-only modeling, image-only modeling, text-to-image, and image-to-text, with a mixture of 70% text-only, 20% image understanding, 9% text-to-image, and 1% image-only. Generation at inference is sequential: the model starts the next chunk from Gaussian noise, integrates the learned vector field from t=0 to t=1 conditioned on the preceding chunks, then appends the generated chunk to the context; because completed chunks do not depend on future chunks, keys and values can be cached. Downstream finetuning (visual question answering, text-to-image) reuses the exact same backbone interface and objective, only changing the task sequence and data. Encoders and decoders — including a separately pretrained six-layer bidirectional text decoder of width 512 and a pretrained image decoder — remain frozen throughout flow pretraining and downstream finetuning.

Why This Matters

The paper argues that neither dominant paradigm simultaneously provides continuous visual states and a common generative process across modalities: fully discrete models tie visual fidelity to the tokenizer, while hybrid models require modality-dependent objectives and sampling procedures. Multimodal Flow's evidence suggests a third option is practical — one objective, one backbone, one sampling procedure — and that a randomly initialized continuous backbone pretrained on mixed multimodal data transfers substantially better to downstream vision-language tasks than one trained on downstream data alone at the same total token budget.

Real-world applications (as implied by the evaluated tasks):

  • Text-to-image generation and image editing pipelines, where compositional prompt following (counting, position, attribute binding, relations) matters — the capabilities measured by GenEval and DPG-Bench.
  • Image captioning and image-conditioned text generation, which the paper evaluates with CIDEr and CLIPScore and where it shows CFG helps.
  • Visual question answering assistants, using the POPE, MMBench, SEEDBench, VQAv2, GQA, and OK-VQA evaluations.
  • Unified multimodal assistants that need one model to both answer questions about images and generate images, served from a single set of weights and a single sampling mechanism.

Industry relevance: The results at 150B pretraining tokens are notable for a from-scratch model at 1.6B parameters, and the paper reports that the same architecture outperforms Transfusion-style and Chameleon-style baselines under matched data, optimization, and parameter budgets. That combination — one backbone for both understanding and generation, with KV caching during sequential generation and reusable frozen codecs — is directly relevant to teams building or serving unified multimodal systems, and the code and model are publicly released at github.com/hustvl/Multimodal-Flow.

Future Directions

  • Longer interleaved sequences: The paper explicitly names extending the framework to longer interleaved multimodal sequences as future work.
  • Video and other structured modalities: Video and other structured modalities are also listed as a promising direction that the authors leave as future work.
  • Scaling beyond the studied range: Only 0.6B, 1.2B, and 1.6B variants are reported, and only the 1.6B model is evaluated as the main benchmark model; whether the trends hold at larger scale and longer token budgets is not reported.
  • Understanding performance gap: MF-1's understanding numbers at 150B tokens trail similarly sized models initialized from pretrained LLMs (for example, MMBench 67.2 versus 74.9 for JanusFlow at 1.3B and 75.5 for Janus-Pro at 1.5B), leaving open how to close that gap within a fully continuous from-scratch recipe.

Target Audience

Researchers and engineers working on unified multimodal foundation models — particularly those evaluating discrete-token, hybrid discrete–continuous, and continuous alternatives; practitioners interested in Flow Matching applied beyond pixels and into text embedding space; and teams building systems that must both understand and generate images with a single backbone and a single objective. Readers without background in diffusion/flow sampling, chunk-causal attention, and classifier-free guidance will find the paper advanced, though the architectural description is largely self-contained.

Authors’ abstract

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The latter requires modality-dependent objectives and sampling procedures. Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains underexplored for multimodal pretraining. Multimodal Flow introduces a unified continuous architecture that integrates multimodal continuous representations with a shared chunk-causal flow backbone. It organizes text blocks and images as ordered continuous hyperchunks, preserving textual token order and visual spatial structure. The backbone learns a single vector field over these hyperchunks through Flow Matching. Joint attention enables cross-modal interaction, while modality-specific feed-forward networks process each modality. The model predicts multiple target chunks in parallel during training and generates hyperchunks sequentially at inference. We instantiate MF-1 and pretrain it on multimodal data. Across 0.6B, 1.2B, and 1.6B scales, continued pretraining consistently improves multimodal modeling. With only 150B pretraining tokens, MF-1 achieves an average score of 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench, and POPE, remaining competitive with unified models trained on substantially more data. Under matched data, optimization, and parameter budgets, Multimodal Flow further outperforms representative hybrid and discrete models. These results establish continuous chunk-based embedding flow modeling as a new fully continuous paradigm for unified multimodal modeling. The related code and model are publicly released at https://github.com/hustvl/Multimodal-Flow.

Read the original paper