Research
LiFT: Loop Flow Transformers
LiFT: Loop Flow Transformers Overview Research area: Generative modeling with diffusion/flow-matching transformers; specifically, looped (weight-shared, recurrent-depth) architectures for class-condit

- arXiv
- 2610.05538
- Published
- 2026-10-04
- Authors
- Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent, Cees G. M. Snoek
AI summary
LiFT: Loop Flow TransformersOverview
Research area: Generative modeling with diffusion/flow-matching transformers; specifically, looped (weight-shared, recurrent-depth) architectures for class-conditional image generation.
Technical level: Advanced. The paper assumes familiarity with flow matching, neural ODEs, Diffusion Transformers, and transformer block design.
Scope: The paper introduces Loop Flow Transformers (LiFT), a training objective that assigns each recurrent pass of a shared Diffusion Transformer core its own point on a straight path from the model's initial velocity estimate to the flow-matching target, and evaluates whether such models keep improving when run at inference depths far beyond their training depth on ImageNet 256×256.
What This Paper Is About
Diffusion Transformers (DiTs) have improved largely by growing parameter count, while looped transformers offer an alternative: repeatedly applying one shared set of layers to increase executed depth at fixed parameter count. The problem is that looped generative models so far only produce coherent images near the depth they were trained at — beyond that, quality usually degrades, and at equal compute a looped DiT has not beaten a dense one. LiFT's goal is to make extra loops at inference time actually improve generation, so a smaller model can match or beat a larger dense DiT.
Key Contributions
- A trajectory objective for recurrent generation. Rather than training every loop toward the complete output, LiFT gives each loop its own target along a straight path from the detached initial estimate
b = sg(u_0)to the flow-matching targetu*, indexed by a continuous depth coordinates. The last loop (s = 1) is trained on the standard flow-matching objective, plus an auxiliary prelude loss weighted byλ. - Useful recurrence far beyond training depth. Without retraining, early exits, or other modifications, an XL/2 checkpoint trained with two loops reduces FID from 17.07 to 9.31 at sixteen loops — eight times its training loop count.
- Better generation with fewer parameters and lower inference cost. On ImageNet 256×256, LiFT-L/2 lowers FID by 3.34 relative to the dense DiT-XL/2 baseline, with approximately 60% fewer parameters, 52% fewer inference FLOPs, and the same training-token budget (also about 32% less training computation: 61.98 versus 91.10 EFLOPs).
- Depth as an axis of inference compute independent of generative time. The loop count
Kand the number of integration stepsTare separated and can both be chosen after training — "Loop in depth, flow in time."
Main Findings
- Training depth is not a ceiling. L/2 R10 improves from FID 20.30 at its two-loop training depth to 10.95 at eight inference loops; XL/2 R12 from 17.07 to 9.31 at sixteen loops; XL/2 R8 from 16.95 at three training loops to 9.50 at thirty-two loops. The largest measured ratio is XL/2 R8 at 10.67× its training depth.
- Single-block cores plateau. Checkpoints with one shared block (B/2 R1) barely improve with extra loops (45.70 at training depth to 45.58 at sixteen loops), which the authors attribute to the limited capacity of a single shared block.
- Larger cores overtake dense DiTs of their scale. L/2 R10 overtakes both dense L/2 and dense XL/2 at three inference loops; XL/2 R12 overtakes dense XL/2 at three; XL/2 R8 and R6 follow at four and eight loops; L/2 R5 draws level with dense L/2 at eight and overtakes it at sixteen.
- A cheaper crossing exists. L/2 R10 at three inference loops beats dense XL/2 (FID 14.64 versus 15.54) with slightly less inference compute (11.43 versus 11.86 TFLOPs) and 60% fewer parameters.
- At its training operating point, LiFT trails a dense model of the same executed depth. L/2 R10 reaches FID 20.30 at two training loops versus 16.94 for dense L/2 at nearly identical cost — consistent with iso-depth scaling findings that a shared recurrence is worth only about half of a unique block. The advantage appears only once test-time compute is reallocated.
- At B/2, dense wins. The best LiFT B/2 result, FID 33.87, does not reach the 32.50 that dense B/2 attains with 25 integration steps at 1.150 TFLOPs; with at most 89M parameters for the same 32.8B training tokens, the cores appear too small for the data budget.
- Cheaper than sampling more, at L/2 and XL/2. LiFT L/2 R10 with four inference loops attains FID 12.25 at 14.79 TFLOPs, whereas dense L/2 needs 16.14 TFLOPs and 100 steps to reach 16.17; LiFT XL/2 R8 with four inference loops reaches FID 11.36 at 15.25 TFLOPs versus 14.69 at 23.72 TFLOPs for dense XL/2 (41% and 56% fewer parameters respectively).
- Balance loops against steps. LiFT L/2 R10 with 25 integration steps and four inference loops achieves FID 13.16 at 7.397 TFLOPs per image, versus 16.94 at 8.069 TFLOPs for dense L/2 with 50 steps — improving on its own training configuration (20.30 at 8.070 TFLOPs) with roughly 41% fewer parameters and 8% less inference computation. But sixteen inference loops with ten steps gives a worse and more expensive result: FID 15.09 at 11.027 TFLOPs.
- The FID surface flattens at larger budgets. With eight inference loops, doubling steps from 50 to 100 leaves FID nearly unchanged (10.95 and 10.96), while extending to thirty-two loops raises FID at both budgets. The best measured setting, FID 10.95 at 28.24 TFLOPs, remains well below the 15.74 dense L/2 reaches with 250 steps at 40.35 TFLOPs.
- Parameter budgets and compute budgets favor different checkpoints. At a target FID of 20, minimizing compute selects LiFT L/2 R10 with ten integration steps and four inference loops (2.96 TFLOPs per image), whereas minimizing parameters selects the smaller LiFT L/2 R5 (176M parameters).
Methodology in Plain English
A flow model generates an image by integrating a learned velocity field from noise to data, and normally you add compute by taking more integration steps — but each step can be no better than the network's velocity estimate, so quality saturates. LiFT instead spends extra compute inside each step by applying a shared transformer core several times.
The architecture splits a DiT into three parts: a prelude that embeds the input once, a shared recurrent core applied K times, and a prediction-only coda (a shared head). Each core application is told where it sits along a continuous depth coordinate s, using a depth embedding injected through adaptive normalization and residual gates; the previous state is normalized and the prelude state reinjected at each pass.
The training trick is what each loop is asked to predict. The model's own first prediction defines an anchor b, detached from gradients. A straight line is drawn from that anchor to the flow-matching target, and the loop at coordinate s is trained to predict the point (1−s)b + s·u*. Intermediate coordinates are drawn randomly from U(0,1) each update and sorted, with endpoints fixed at 0 and 1; the last readout therefore always does ordinary flow matching. An extra prelude regression loss (λ = 0.01) keeps the anchor from drifting.
At inference, the loop budget can change freely: coordinates become a uniform grid s_k = k/K_inf, and only the final readout is fed to a Euler sampler. Because progress along the path is continuous, the loop count only sets how finely the path is traversed — just as the number of integration steps sets how finely generative time is traversed.
Experiments use class-conditional ImageNet-1k at 256×256 with SD-VAE latents of shape 4×32×32. All models are trained on a fixed budget of D = 3.2768 × 10^10 latent-patch tokens (500,000 updates at batch size 256) with AdamW at learning rate 10⁻⁴. Within each scale, the executed transformer depth 4 + K_train·L_R is held equal to the dense reference: 12 blocks for B/2, 24 for L/2, 28 for XL/2. Evaluation samples 50,000 images per setting with EMA weights and Euler integration without classifier-free guidance, reporting FID, sFID, and IS via the official ADM evaluator (FID primary).
Why This Matters
Impact on research. The paper reframes recurrent depth in generative models as a supervision problem rather than just an architecture problem. Its key claim is that asking every intermediate prediction for the complete output gives no stage a distinct role, whereas assigning each loop a fraction of the correction makes extra loops useful beyond training depth. It provides a concrete counterpoint to prior looped diffusion results (Goyal et al., 2026; Chai, 2026), where loops beyond training depth hurt FID or failed to beat dense models at equal compute, and it connects to iso-depth scaling work (Schwethelm et al., 2026) and depth-conditioned looped transformers (Xu and Sato, 2025; LoopFormer, Jeddi et al., 2026).
Real-world applications (envisioned, not reported as evaluated in the paper):
- Image generation services with fixed memory budgets, where the same checkpoint serves both a cheap low-loop setting and a higher-quality high-loop setting without storing a second set of weights.
- On-device or edge generation, where the parameter count and memory footprint matter more than latency, since the paper reports every quality target reached with fewer parameters than a dense DiT.
- Latency-vs-quality tuning in deployment, because loops (
K_inf) and integration steps (T) can be chosen after training to hit a compute budget. - Any flow- or diffusion-transformer pipeline where test-time compute is elastic, using the model-selection analysis (least compute or fewest parameters for a target FID) as a guide.
Industry relevance. The work targets the practical trade-off that dominates generative-model deployment: quality per unit of memory and per unit of inference compute. LiFT reports 60% fewer parameters, 52% fewer inference FLOPs, and about 32% less training computation for its headline comparison against dense DiT-XL/2, with no classifier-free guidance used at sampling.
Future Directions
- Understanding the capacity threshold. The B/2 scale fails to overtake dense B/2 (best LiFT FID 33.87 versus 32.50), which the authors attribute to cores that are too small — the exact condition under which recurrence becomes worthwhile is not resolved.
- Choosing the loop/step split automatically. The paper shows the best allocation varies with the starting configuration (four loops at 25 steps beats sixteen loops at 10 steps for the same checkpoint), but gives no rule for predicting it.
- Extending beyond ImageNet 256×256 class-conditional images. Video is mentioned only through prior work (Goyal et al., 2026), where quality peaked at 1.5× training depth; other resolutions, modalities, and text conditioning are not evaluated.
- Comparing trajectory supervision against alternatives. The paper positions itself against intermediate supervision that regresses the complete output, and against MeanFlow-trained LoopDiT (Chai, 2026), but a systematic comparison of objectives under matched budgets is not reported.
Target Audience
Researchers and engineers working on diffusion and flow-matching generative models, especially those interested in efficient architectures, weight sharing, test-time compute scaling, or reducing parameter count without losing sample quality. It is most useful to readers already comfortable with flow matching, DiT architectures, and FID-based evaluation; practitioners focused on deployment trade-offs will also find the compute-versus-parameter analyses directly actionable.
Authors’ abstract
We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate, a trained model can loop far beyond its training depth with no retraining, early exits, or other modifications. In our experiments, these longer rollouts improve generation, so inference computation can grow without adding parameters. On ImageNet at 256x256, LiFT-L/2 achieves an FID 3.34 points lower than our dense DiT-XL/2 baseline while using approximately 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.