Research
MEND: RL For Flow Models via Proximal Velocity Matching
Overview Research area: Reinforcement learning and reward post-training for flow-based (rectified-flow) text-to-image generative models. Technical level: Advanced. The paper assumes familiarity with f

- arXiv
- 2610.05954
- Published
- 2026-10-05
- Authors
- Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
AI summary
Overview
Research area: Reinforcement learning and reward post-training for flow-based (rectified-flow) text-to-image generative models.
Technical level: Advanced. The paper assumes familiarity with flow matching, policy-gradient RL for diffusion models, LoRA adapters, and proximal-point optimization; the appendices contain formal propositions with proofs.
One-sentence scope: MEND is a reward post-training method that gives a flow model a new velocity target for a sample only when the capped reward gained by moving the sample exceeds a quadratic price on how far it moves.
What This Paper Is About
Existing ways to post-train image generators against a learned reward either reweight the model's own samples under a KL penalty or a frozen reference (Flow-GRPO, DiffusionNFT), or backpropagate the reward and move every sample without checking whether the move is worth its size (ReFL, DRaFT). The paper argues that neither approach makes a per-sample decision about whether a change should happen at all, which can suppress hard modes (reweighting) or push samples into low-density regions (unchecked ascent). MEND's goal is to make that per-sample decision explicit, so reward improves quickly without eroding what the base model already does well.
Key Contributions
-
Insight and method. The authors introduce MEND, which caps rewards within each prompt group so well-scoring samples receive no move, proposes a few moves along the normalized reward gradient for samples below the cap, and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets with no KL term, frozen reference model, or advantage weights.
-
Theory. Proposition 1 shows every accepted target improves capped reward by more than its quadratic price, with a displacement bound that shrinks to zero as the sample approaches the cap. Proposition 2 shows the regression gradient is reward backpropagation restricted to accepted samples and scaled by their selected steps.
-
Performance. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43.
-
Generality. The method is described as backbone-agnostic and improves SD3.5-M, SD3-M, and the distilled Z-Image-Turbo at the same update budget, requiring only a differentiable reward.
Main Findings
-
Fewer updates, comparable or better reward. With 100 updates, MEND beats Flow-GRPO on five of six evaluators at the same DreamSim distance to base images (0.313). Aesthetic score is the single evaluator where Flow-GRPO is higher (5.90 versus 5.88).
-
Equal-budget wins over ReFL and DiffusionNFT. Under the equal-budget protocol, MEND reaches PickScore 24.03 at update 100, versus 23.92 for ReFL and 23.43 for DiffusionNFT, and is ahead of both baselines at every evaluated update. On 64 unseen Pick-a-Pic prompts it is above DiffusionNFT at every evaluated update and by update 25 surpasses the levels of Flow-GRPO and DiffusionNFT, which use about 4k and 1.7k updates.
-
Three-reward run beats a five-reward baseline on its own rewards. A 300-update run trained on PickScore, HPSv2.1 and CLIPScore surpasses the five-reward DiffusionNFT model on those three evaluators, but is lower on HPSv3, ImageReward and aesthetic score, which it does not train on. Table 2 lists that run's base distance as 0.488 versus 0.538 for the five-reward DiffusionNFT.
-
Works across training rewards. Under the equal-budget protocol, MEND surpasses ReFL and DiffusionNFT on each of four training rewards (PickScore, ImageReward, HPSv2.1, CLIPScore) at update 100 and at every evaluated update. CLIPScore and ImageReward training yield the lowest held-out PickScore (22.27 and 22.33), and three runs show held-out declines before update 100.
-
Works across backbones. On SD3-M, PickScore rises from 20.36 to 23.70, above Linear-DPO at 20.96, with held-out HPSv2.1 rising from 0.215 to 0.298. On the distilled Z-Image-Turbo (nine steps, 1024 pixels), PickScore training raises that score from 22.86 to 24.10 and held-out HPSv2.1 from 0.295 to 0.315; HPSv2.1 training reaches 0.357 with held-out PickScore 23.27.
-
Reward keeps rising while drift stays nearly flat. Between updates 25 and 100, DrawBench PickScore rises from 23.36 to 23.70 while base distance moves only from 0.294 to 0.313.
-
Cost and stability. Training costs 10.0 GPU-hours on 3 GB200 GPUs, excluding evaluation. A second launch of the same configuration differs by 0.01 PickScore at update 50.
-
The cap and price trade reward for diversity. MEND accepts moves for 34% of seeds per round, reaching PickScore 23.10 at diversity 0.297. Removing the cap or the price raises the acceptance share to 45% and 78% and lowers diversity to 0.287 and 0.291. Moving every seed below the cap (81%) gives the highest full-set PickScore, 23.45, and lowest diversity, 0.260. Against that arm, the full rule keeps 0.037 more diversity for 0.35 PickScore.
-
Target-checking is the distinctive component. Table 1 contrasts MEND (no KL or frozen reference, reward weights, reward gradient, target check, 100 updates) against Flow-GRPO (KL/frozen reference, reward weights, about 4k updates), ReFL (gradient, no check, 100 updates) and DiffusionNFT (KL/frozen reference, reward weights, 1.7k updates).
-
Other held-out checks. MEND scores 0.62 on GenEval against 0.32 for its base at guidance 1; at guidance 4.5 it matches the Flow-GRPO PickScore adapter on GenEval (0.74 against 0.74) and is below it on OCR (0.53 against 0.68). At equal diversity it is at least 0.35 above DiffusionNFT under the same protocol.
Methodology in Plain English
A rectified flow interpolates between a clean latent and Gaussian noise. At a given timestep, the model's velocity gives a clean prediction of the latent. MEND keeps two copies of the base model: a trained adapter and a lagged behavior adapter.
Each training round has three parts. First, the behavior adapter generates G deterministic 10-step trajectories per prompt and the endpoint of each is scored. Within each prompt group, rewards are capped at the 0.75 quantile, subject to a global floor that rises slowly across rounds. Samples at or above the cap get no move. Second, each sample below the cap whose reward gradient is non-zero receives three proposed moves along its normalized reward gradient, at step lengths 0.1, 0.2 and 0.4; each proposal is decoded and scored once. The verdict subtracts a quadratic displacement price (distance squared divided by twice a price parameter tau) from the capped reward and keeps the best candidate, with the unchanged sample always among the candidates and ties favoring it. Because the reward is evaluated on decoded candidates rather than predicted from a first-order approximation, the decision rests on the reward actually attained. Third, the model regresses its velocity at a stored rollout state nearest t = 0.278 onto the behavior velocity shifted by the accepted displacement, with the endpoint itself never used as the target. Kept samples are anchored to the behavior prediction by an additional term weighted at 10. The behavior adapter follows the trained adapter by EMA with coefficient min(0.001u, 0.5), so the anchor tracks training instead of being frozen. Each round takes one AdamW step at learning rate 3×10⁻⁴ with epsilon 10⁻¹², and tau is adjusted between rounds to target a 30% to 60% acceptance rate among proposed samples. Evaluation uses a separate EMA of the trained parameters.
The evaluation protocol trains SD3.5-M on PickScore with 48 Pick-a-Pic prompts times 24 images per update, rank-32 LoRA, 10-step rollouts at 512 pixels, and 100 updates, then evaluates on DrawBench at 200 prompts times 5 seeds with 40 Euler steps, scoring PickScore, HPSv2.1, HPSv3, ImageReward, CLIPScore and aesthetic quality, with DreamSim distance to the base image as the fidelity measure.
Why This Matters
For research, MEND reframes reward post-training as per-sample target selection with an explicit cost, rather than distribution-level regularization or uniform reweighting. That yields two formal guarantees about accepted targets and shows the regression gradient is reward backpropagation restricted to accepted samples, which connects the method to earlier reward-gradient approaches while sharpening them. It also removes three pieces of machinery commonly needed in this setting — KL penalties, a frozen reference model, and advantage weights — while remaining backbone-agnostic.
Real-world applications:
- Consumer image generation and editing tools that need outputs aligned to human preference without visual drift from the base model, since the method controls DreamSim distance to the base image.
- Advertising and marketing asset generation, where prompt fidelity (counts, attributes, spatial relations) matters and where MEND is shown correcting base-model prompt errors from the same noise.
- Games, film and design workflows that post-train or customize a generator for a specific house style or reward model at low compute.
- Model providers serving distilled generators, since the method improves Z-Image-Turbo, a distilled model sampled in nine steps at 1024 pixels.
Industry relevance: a full SD3.5-M run costs 10.0 GPU-hours on three GB200 GPUs excluding evaluation, which the authors state makes reward post-training cheap enough to repeat for each new reward or backbone. That is a practical argument for per-release reward tuning rather than one-off large runs.
Future Directions
- Reducing reliance on a differentiable reward, since the method's guarantees and mechanics require one.
- Understanding and controlling reward over-optimization under longer training, which the limitations section names and which appendices E.3 and F.4 examine.
- Addressing the held-out declines seen in some single-reward runs within the 100-update budget, and the lower held-out PickScore from CLIPScore and ImageReward training (22.27 and 22.33).
- Closing the gap to Flow-GRPO on OCR (0.53 against 0.68 at guidance 4.5) and on aesthetic score, and examining the higher saturation that many MEND examples show.
- Extending the same per-sample acceptance rule to further backbones, rewards and sampling settings beyond those tested.
Target Audience
Researchers and engineers working on reward alignment and post-training for diffusion and flow-based generative models, particularly those with limited compute who cannot afford thousands of RL updates or KL-regularized pipelines. It is also relevant to practitioners who need to customize a text-to-image backbone against a proprietary reward model, and to readers interested in proximal-point arguments applied to generative model fine-tuning. The appendix theory and ablation tables would be most useful to readers already comfortable with policy-gradient and weighted-regression baselines.
Authors’ abstract
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.