Skip to content
AI.info

Research

SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

Overview Research area: Computer vision, specifically generative video — high-resolution video generation, diffusion-model refinement, and inference acceleration. Technical level: Intermediate. The pa

SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
arXiv
2609.37969
Published
2026-09-29
Authors
Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen, Haopeng Li, Jincheng Yu, Duomin Wang, Ruihua Zhang, Lei Zhu, Song Han, Enze Xie

AI summary

Overview

Research area: Computer vision, specifically generative video — high-resolution video generation, diffusion-model refinement, and inference acceleration.

Technical level: Intermediate. The paper is readable without deep mathematical background, but familiarity with diffusion/flow-matching sampling, denoising steps (NFEs), and distillation concepts helps with the method sections.

Scope: This paper introduces SoL-Refiner, a video refiner that converts low-resolution generator outputs into high-resolution video (up to 4K) in a single denoising step, trained through a three-stage recipe and benchmarked against existing refiners on a newly introduced 150-video benchmark called Refiner-Bench.

What This Paper Is About

High-resolution video generation is expensive because the cost grows rapidly with the number of spatiotemporal tokens, and multiplying those tokens across many denoising steps makes direct high-resolution sampling impractical. A common workaround is a two-stage pipeline: generate a low-resolution video first, then apply a refiner — but conventional refiners need multiple target-resolution denoising steps, creating a second sampling bottleneck. The goal of this work is a refiner that does the job in a single step, works on outputs from several different base generators, and can be deployed quickly enough to be practical.

Key Contributions

  1. A one-step video refiner. SoL-Refiner upsamples and refines low-resolution outputs from multiple base generators in a single denoising step, without modifying or retraining the base models. It is initialized from a pretrained LTX-2.3 checkpoint.
  2. A three-stage training recipe. High-resolution continual training learns the refinement mapping, reinforcement learning (RL) post-training with frame-based reward models improves perceptual quality, and staged distillation compresses the multi-step refiner into a one-step model.
  3. Refiner-Bench. A 150-video refinement benchmark built from the outputs of different video generators (50 videos each from WAN 2.1 1.3B, SANA-Video 2B, and LTX-2 Stage 1), evaluated under a shared-input protocol so that refiners are compared on identical inputs rather than on the choice of upstream generator.
  4. An acceleration stack. A tiny autoencoder (TAE) replaces the full video VAE, and Sol-Engine provides kernel fusion and sparse attention, together giving an 8.91× speedup in refinement latency over the three-step LTX-2.3 Refiner in the paper's 2K latency setting.

Main Findings

  • One-step refinement beats all evaluated external refiners at 2K. Under the shared-input protocol with all 150 videos resized to 1024×576, the one-step SoL-Refiner reaches a VBench AVG of 0.81048 and a UniPercept AVG of 60.4150, outperforming LingBot Stage-2 Refiner, LTX-2.3 Refiner, LTX-2.0 Refiner, and SEEDVR2 on VBench AQ, IQ, and AVG and on UniPercept AVG, including SEEDVR2 under the same one-step budget.
  • The 23-step variant scores higher still. SoL-Refiner (Multi-Step, 23 steps) reaches VBench AVG 0.81691 and UniPercept AVG 61.3170, versus 0.81048 and 60.4150 for the one-step model.
  • Improvement over the direct baseline grows with resolution. Compared with the official three-step LTX-2.3 Refiner at 720p, 2K, and 4K, SoL-Refiner scores higher on both VBench and UniPercept. At 4K (3840×2176) it improves VBench AVG by 3.86% and UniPercept AVG by 22.79% over that baseline.
  • Refinement latency drops from 57.461 to 6.447 seconds. In the 2K latency setting, adding one-step distillation, TAE, and Sol-Engine in sequence gives adjacent speedups of 1.82×, 2.37×, and 2.06×, for a final 8.91× speedup over the three-step LTX-2.3 Refiner.
  • The two-stage pipeline is much faster than direct high-resolution generation. For a 5-second, 1344×768, 24-fps video, the published full-resolution 50-step SGLang MiniMax H3 baseline takes 152.3 s on one GB200, whereas distilled four-step 896×512 MiniMax H3 generation followed by one-step SoL-Refiner takes 5.64 s using one GPU per stage — a 27.03× speedup, split into 4.06 s for H3 generation and 1.56 s for the refiner.
  • Acceleration generalizes across base generators. Running WAN-5B, WAN-1.3B, and Cosmos-Nano at low resolution and then refining with the one-step SoL-Refiner reduces latency by 54.7%, 71.1%, and 64.4% respectively, while improving mean VBench and UniPercept over direct high-resolution generation. For example, WAN-5B direct high-resolution runs at 80.15 VBench / 54.18 UniPercept in 87.70 s, while low-resolution plus SoL-Refiner gives 81.22 / 54.49 in 39.75 s with 6.16 s of refiner overhead.
  • RL post-training is the highest-scoring stage of the recipe. In the training-recipe ablation, Stage 2 (post-training, multi-step) reaches VBench AVG 0.81691 and UniPercept AVG 61.3170, above Stage 1 continual training (0.80888 / 56.6620). After distillation, the one-step model remains above the continual-training baseline on both averages.
  • Both reward design choices help. Removing the regularization recipe lowers VBench AVG to 0.81046, and using only HPSv3++ instead of multiple reward models gives 0.81415, both below the full configuration's 0.81691.
  • Sequential RL-then-DMD beats joint optimization. DMD-R (jointly optimizing RL and DMD objectives) scores VBench AVG 0.80098, versus 0.81048 for RL post-training followed by DMD, so RL then DMD is used in all other experiments.
  • Sol-Engine speeds up multi-step refinement too. Across four resolution and frame-count settings, Sol-Engine gives 2.38× to 3.26× speedups over the multi-step base pipeline. As an example, at 2K with 81 frames, latency drops from 127.7 s to 53.7 s on one H100.
  • Deployment on a single DGX Spark. On one NVIDIA DGX Spark (GB10), using SoL-Refiner in Stage 2 of a SoL-H3 pipeline (H3 generation at 384p, refinement to 768p) with Stage 1 unchanged reduces end-to-end latency from 56.17 to 39.67 seconds (29.4%), a 9.43× speedup over direct four-step H3 generation at 768p against a 374-second full-resolution four-step baseline. Stage 1 takes 21.60 seconds in both configurations; Stage 2 falls from 34.57 to 18.07 seconds.
  • A stated limitation. On Refiner-Bench, the 23-step model achieves higher VBench and UniPercept averages than the one-step model, so the acceleration comes with a measurable quality trade-off on these metrics.

Methodology in Plain English

The refiner is trained to map a low-quality video to a high-quality one rather than to generate from scratch. The researchers start from a pretrained high-capacity video diffusion model (an LTX-2.3 checkpoint) and train it in three stages.

In the first stage, the model learns the refinement mapping from paired low-fidelity and high-fidelity videos that share

Authors’ abstract

High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at $3840\!\times\!2176$ it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an $8.91\times$ speedup in refinement latency over the same baseline in our 2K latency setting.

Read the original paper