Research
FastFlow: Accelerating The Generative Flow Matching Models with Bandit Inference
Overview Research area: Computer vision / generative modeling, specifically inference acceleration for flow-matching (FM) generative models. Technical level: Intermediate. The paper assumes familiarit

- arXiv
- 2602.11105
- Published
- 2026-02-11
- Authors
- Divya Jyoti Bajpai, Dhruv Bhardwaj, Soumya Roy, Tejas Duseja, Harsh Agarwal, Aashay Sandansing, Manjesh Kumar Hanawal
AI summary
Overview
- Research area: Computer vision / generative modeling, specifically inference acceleration for flow-matching (FM) generative models.
- Technical level: Intermediate. The paper assumes familiarity with flow-matching ODEs, numerical solvers (Euler method), and multi-armed bandit (MAB) theory, though the core intuition is explained clearly.
- Scope: The paper introduces FastFlow, a training-free, plug-and-play inference framework that uses finite-difference velocity extrapolation and multi-armed bandit decisions to skip redundant denoising steps in flow-matching models for image generation, image editing, and video generation.
What This Paper Is About
Flow-matching models produce high-quality images and videos but are slow because they generate samples through many sequential denoising steps, each requiring a full neural network evaluation. Existing accelerators such as distillation, trajectory truncation, and consistency training are static, need retraining, and often do not generalize across tasks. FastFlow addresses this by adaptively deciding, at inference time and per sample, which denoising steps can be safely approximated instead of fully computed.
Key Contributions
- An adaptive skipping framework. FastFlow skips redundant denoising steps using a simple Euler solver for the flow and a first-order Taylor series expansion for velocities, framing the speed-versus-error trade-off as a multi-armed bandit (MAB) objective so the model dynamically learns when full computation is necessary.
- A theoretical error bound. The paper establishes a bound (Theorem 3.1) on the deviation of the final state from the approximated trajectory relative to the full-model trajectory.
- A model-agnostic, training-free design. The framework requires no retraining and no auxiliary networks, and integrates into existing flow-matching pipelines.
- Broad empirical validation. Experiments across image generation, video generation, and image editing demonstrate more than 2.6× speedup while maintaining generation quality.
Main Findings
- Speedup versus quality: On text-to-image generation (Table 1), FastFlow-50 reports an Overall score of 0.78, CLIPIQA of 0.83, 2.65× speedup, and 13.7 s latency, compared with the Full 50-step model's Overall 0.78, CLIPIQA 0.85, 1.00× speedup, and 36.2 s latency. FastFlow-25 reports Overall 0.77, CLIPIQA 0.80, 4.54× speedup, and 8.6 s latency; FastFlow-10 reports Overall 0.72, CLIPIQA 0.73, 7.14× speedup, and 5.5 s latency.
- Comparison with static baselines (Table 1): InstaFlow reports Overall 0.33, CLIPIQA 0.74, 50.0× speedup, and 1.5 s latency; PerFlow reports Overall 0.58, CLIPIQA 0.80, 5.00× speedup, and 8.2 s latency; TeaCache reports Overall 0.76, CLIPIQA 0.80, 1.85× speedup, and 20.6 s latency.
- Second evaluation table (Table 2): The Full 50-step model reports Overall 0.65, CLIPIQA 0.84, 1.00× speedup, 33.8 s latency; TeaCache-50 reports Overall 0.64, CLIPIQA 0.80, 1.91× speedup, 18.3 s latency; FlowFast (as labeled in the paper) 50 reports Overall 0.64, CLIPIQA 0.82, 2.57× speedup, 13.9 s latency; FlowFast 25 reports Overall 0.63, CLIPIQA 0.79, 4.21× speedup, 8.5 s latency; FlowFast 10 reports Overall 0.55, CLIPIQA 0.57, 7.59× speedup, 5.2 s latency.
- Theoretical guarantee: Theorem 3.1 states that under smoothness assumptions and uniform step size Δt = 1/T, the cumulative final-state error after T steps satisfies e_T = O(|S| / T³), where S is the set of skipped steps. The error therefore grows linearly with the number of skipped steps.
- Motivating observation (Figure 4): Analysis of the L1-relative error between consecutive velocity predictions in the BAGEL model reveals a three-phase pattern — the model first establishes the coarse flow, then performs subtle refinements during intermediate steps, and finally adjusts again in later steps. The paper argues these intermediate refinements are small but not negligible.
- Reward design: The reward for an action α is r(α) = μ · α − ℓ(v̂, v), where μ balances efficiency against accuracy and ℓ is a discrepancy measure such as mean-squared error. The paper sets μ = max_t MSE(v̂_t, v_t) / total steps, estimated from the first full generation pass.
- Bandit convergence (Figure 8): Cumulative regret across bandits at different denoising positions begins to flatten between 50 and 100 samples, indicating rapid identification of a near-optimal skipping strategy.
- Image editing (Figure 2): Evaluated on the GEdit dataset with BAGEL and FLUX models, using GPT-4.1 as an automatic judge scoring semantic consistency (G_SC), perceptual quality (G_PQ), and overall score (G_O). The paper reports the highest speedup among baselines while preserving or improving edit quality. Specific numerical values are not reported in the truncated content.
- Video generation (Figure 3): Evaluated on VBench with the HunyuanVideo model, additionally reporting the no-reference BRISQUE metric for frame quality. The paper reports sharper frames and more coherent temporal evolution with substantial acceleration; specific VBench and BRISQUE values are not reported in the truncated content.
- Adaptive skip patterns (Figure 6): When velocity fluctuations are high, FastFlow chooses shorter skips; in smoother intermediate regions it shifts to longer skips; as fluctuations re-emerge toward later steps it reduces skip length again.
- Stated limitation: The speedup may not be observed in the initial steps due to the inherent exploration phase of the multi-armed bandit.
Methodology in Plain English
Flow-matching models learn a velocity field that moves samples along a path from noise to data, and inference means repeatedly asking the neural network for the velocity at each timestep. FastFlow observes that many of these steps only make minor adjustments, so it approximates them instead of calling the network.
The approximation works like this: rather than reusing the previous velocity unchanged (as earlier caching methods do), FastFlow uses a finite-difference estimate of how the velocity has been changing between prior predictions, and extrapolates the velocity forward with a first-order Taylor expansion. This lets the trajectory advance several steps at essentially zero compute cost. A full model evaluation is then performed at the point where the excursion ends, which serves as a check on how good the approximation was.
The crucial decision — how many steps to skip before the next real model call — is treated as a multi-armed bandit problem, with a separate bandit instantiated at each timestep. Each arm corresponds to a skip length. The bandit picks an arm using an upper-confidence-bound strategy (with exploration constant γ set to 2.0 alongside μ = 2.0), receives a reward that credits the number of skipped steps but penalizes the mismatch between the extrapolated velocity and the actual velocity returned by the model, and updates its statistics. This makes the skipping policy adaptive per sample: easy inputs are skipped aggressively, harder ones trigger more full evaluations. The method requires no retraining and no extra networks.
Evaluation spans the GenEval benchmark (553 prompts for compositional text-to-image reasoning), the GEdit benchmark (606 real-world English editing instructions), and a VBench subset (80 prompts, sampled as 5 from each of 16 dimensions). Models tested include BAGEL, Flux-Kontext, and PeRFlow for image generation; BAGEL, Flux-Kontext, and Step-1X-Edit for image editing; and HunyuanVideo for video generation. All experiments ran on a single NVIDIA A100 GPU.
Why This Matters
- Impact on research: FastFlow offers an alternative to retraining-based acceleration (distillation, consistency training) by showing that adaptive, training-free inference-time skipping can deliver competitive speedups. It also supplies a formal error bound for trajectory approximation, which is uncommon among caching-style acceleration methods.
- Real-world applications:
- Real-time or near-real-time text-to-image generation on consumer or resource-constrained hardware.
- Interactive image editing tools that need low-latency response to instructions.
- Video synthesis pipelines where long durations and high resolutions make per-step cost prohibitive.
- Deployment in compute-limited production environments where retraining or maintaining additional distilled checkpoints is impractical.
- Industry relevance: The work is a collaboration involving Amazon and IIT Bombay affiliations, and it is funded in part by the Amazon IIT-Bombay AI-ML Initiative (AIAIMLI), alongside support from the Prime Minister's Research Fellowship, the Telecom Centre of Excellence (TCOE), the Department of Telecommunication (DoT), and the Ministry of Electronics and Information Technology (MeitY). Because FastFlow is plug-and-play and model-agnostic, it targets the practical constraint that generative pipelines must be accelerated without replacing already-deployed models.
Future Directions
- Reducing the exploration cost at early steps. The authors explicitly note that speedup may not appear in initial steps because of the bandit's exploration phase; shortening or warm-starting this phase is a natural next step.
- Combining FastFlow with other acceleration families. The paper positions FastFlow against distillation, truncation, and consistency methods without testing combinations; layering on top of a distilled or few-step solver is an open question.
- Extending beyond current benchmarks and modalities. Evaluation covers GenEval, GEdit, and an 80-prompt VBench subset; broader video benchmarks, longer durations, and additional model families are untested.
- Generalizing the arm-set and reward design. The arm set is fixed across models and tasks and only updated when the generation horizon changes, and μ is derived from a first full pass; whether more adaptive arm or reward formulations improve the speed-fidelity frontier remains unexplored.
Target Audience
Researchers and practitioners working on generative modeling efficiency, particularly those deploying flow-matching and diffusion-based image or video systems under latency or compute constraints. It is also relevant to readers interested in bandit-based online decision-making applied to systems problems, and to engineers seeking a training-free accelerator that can be dropped into existing pipelines without modifying model weights.
Authors’ abstract
Flow-matching models deliver state-of-the-art fidelity in image and video generation, but the inherent sequential denoising process renders them slower. Existing acceleration methods like distillation, trajectory truncation, and consistency approaches are static, require retraining, and often fail to generalize across tasks. We propose FastFlow, a plug-and-play adaptive inference framework that accelerates generation in flow matching models. FastFlow identifies denoising steps that produce only minor adjustments to the denoising path and approximates them without using the full neural network models used for velocity predictions. The approximation utilizes finite-difference velocity estimates from prior predictions to efficiently extrapolate future states, enabling faster advancements along the denoising path at zero compute cost. This enables skipping computation at intermediary steps. We model the decision of how many steps to safely skip before requiring a full model computation as a multi-armed bandit problem. The bandit learns the optimal skips to balance speed with performance. FastFlow integrates seamlessly with existing pipelines and generalizes across image generation, video generation, and editing tasks. Experiments demonstrate a speedup of over 2.6x while maintaining high-quality outputs. The source code for this work can be found at https://github.com/Div290/FastFlow.