High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at 3840×2176 it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an 8.91× speedup in refinement latency over the same baseline in our 2K latency setting.
Figures & tables
Figure 1 : Overview of SoL-Refiner. (a) SoL-Refiner improves video resolution and visual quality while repairing artifacts in generated videos. (b) For a 5-second, 1344×768 , 24-fps video, the published full-resolution 50-step SGLang MiniMax H3 baseline [ 1 ] takes 152.3 s on one GB200, whereas distilled four-step 896×512 MiniMax H3 generation followed by the one-step SoL-Refiner takes 5.64 s using one GPU per stage ( 27.03× ). The two-stage pipeline spends 4.06 s in H3 generation and 1.56 s in Refiner service. (c) Qualitative examples show how SoL-Refiner improves low-resolution videos generated by MiniMax H3.
Figure 2 : Training and inference pipeline of SoL-Refiner. Left, paired LQ/HQ videos and text/reference conditioning support continual training, RL post-training, and DMD-GAN distillation. Middle, the insets detail the truncated- σ flow path, reward feedback after a no-gradient rollout, and the DMD gradient and DiT discriminator that compress a multi-step teacher into a one-step student. Right, a tiny autoencoder (TAE) [ 14 ] replaces the full video variational autoencoder (VAE) to reduce encoding and decoding cost; latent upsampling and noise initialize the one-step refiner, and Sol Video Inference Engine (Sol-Engine) [ 15 ] executes one refiner NFE.
Method
Eval. Res.
# Steps
VBench ↑
Δ↑
SC
BC
MS
DD
AQ
IQ
AVG
Resized Input
1024×576
–
0.91805
0.95261
0.98354
0.70000
0.60621
0.63158
0.79866
0.00000
LingBot Stage-2 Refiner
1920×1088
8
0.91715
0.95354
0.98398
0.74000
0.57762
0.63792
0.80170
0.00304
LTX-2.3 Refiner
2048×1152
3
0.93004
0.95354
0.98873
0.70000
0.59541
0.65553
0.80388
0.00522
LTX-2.0 Refiner
2048×1152
3
0.93036
0.95909
0.98656
0.72000
0.60679
0.64497
0.80796
0.00930
SEEDVR2
2048×1152
1
0.91804
0.94783
0.97723
0.71333
0.60129
0.66337
0.80351
0.00485
Table 1 : VBench results on Refiner-Bench under a shared-input protocol. All methods receive the same 1024×576 resized inputs. Resolution denotes the evaluated video resolution. SC, BC, MS, DD, AQ, and IQ denote subject consistency, background consistency, motion smoothness, dynamic degree, aesthetic quality, and imaging quality. Δ is the AVG difference from the resized input. Best results are in bold ; second-best results are underlined .
Method
Eval. Res.
# Steps
UniPercept ↑
IAA
IQA
ISTA
AVG
Resized Input
1024×576
–
59.3540
60.8943
42.0890
54.1124
LingBot Stage-2 Refiner
1920×1088
8
58.9189
58.0827
42.1403
53.0473
LTX-2.3 Refiner
2048×1152
3
61.2239
61.9438
45.2052
56.1243
LTX-2.0 Refiner
2048×1152
3
60.5074
60.4902
42.3823
54.4600
SEEDVR2
2048×1152
1
62.5692
64.4827
43.4057
56.8192
Table 2 : UniPercept results on Refiner-Bench under a shared-input protocol. All methods receive the same 1024×576 resized inputs. Resolution denotes the evaluated video resolution. Best results are in bold ; second-best results are underlined .
Figure 3 : Performance on Refiner-Bench across output resolutions. We compare the three-step LTX-2.3 Refiner with the one-step SoL-Refiner using VBench AVG (left) and UniPercept AVG (right). Red Δ labels report the relative improvement over LTX-2.3 Refiner. VBench scores are shown on a 0–100 scale. Both vertical axes use truncated ranges for readability. Higher is better.
Figure 4 : Quality and acceleration across base generators. (a,b) VBench and UniPercept averages for direct high-resolution generation, low-resolution generation, and low-resolution generation followed by SoL-Refiner. (c) End-to-end latency on one H100 GPU, split into base generation and refinement; speedups are relative to direct high-resolution generation. WAN uses 81 frames, 50 generation steps, and 832×480→1280×704 refinement; WAN-1.3B’s direct baseline is 1280×720 . Cosmos-Nano uses 189 frames, 35 generation steps, and 480p →1280×720 refinement.
Training Stage
VBench ↑
UniPercept ↑
SC
BC
MS
DD
AQ
IQ
AVG
IAA
IQA
ISTA
AVG
Stage 1: Continual (Multi-Step)
0.92756
0.95361
0.98201
0.72000
0.60526
0.66484
0.80888
63.0807
64.5693
42.3359
56.6620
Stage 2: Post-Training (Multi-Step)
0.92988
0.94935
0.98027
0.72000
0.62344
0.69851
0.81691
66.7284
72.3977
44.8249
61.3170
Stage 3: DMD-GAN (One-Step)
0.92191
0.94591
0.98054
0.71333
0.60921
0.69196
0.81048
65.7322
70.6682
44.8445
60.4150
Table 3 : Ablation of the three-stage training recipe on Refiner-Bench. We report VBench and UniPercept scores after each stage. Best results are in bold ; second-best results are underlined .
Configuration
VBench ↑
SC
BC
MS
DD
AQ
IQ
AVG
w/o regularization recipe
0.94438
0.95297
0.98074
0.62667
0.64757
0.71044
0.81046
w/o multiple reward models
0.93553
0.94765
0.98443
0.70000
0.62616
0.69111
0.81415
Ours
0.92988
0.94935
0.98027
0.72000
0.62344
0.69851
0.81691
Table 4 : Reward-learning ablation on Refiner-Bench. We report VBench scores. The “w/o multiple reward models” variant uses only HPSv3++. Best results are in bold ; second-best results are underlined .
Figure 5 : Qualitative effect of RL post-training. The rows show the low-quality input, the refiner without RL post-training, and the same refiner with RL post-training. Four aligned frames span 0.0 to 5.0 seconds. Each complete frame is followed by a magnified crop from the red box. In this selected example, RL post-training produces cleaner facial contours and sharper eyes, mouth, and jewelry across time.
Configuration
VBench ↑
SC
BC
MS
DD
AQ
IQ
AVG
DMD-R
0.91385
0.93886
0.97779
0.66000
0.61929
0.69608
0.80098
RL then DMD
0.92191
0.94591
0.98054
0.71333
0.60921
0.69196
0.81048
Table 5 : Distillation ablation on Refiner-Bench. We compare DMD-R with RL followed by DMD using VBench. Best results are in bold ; second-best results are underlined .
Figure 6 : Latency analysis. Panel (a) compares Base (Multi-Step) with the same pipeline using Sol-Engine on one H100 GPU. The four groups report 2048×1152 and 3840×2176 outputs with 81 or 161 frames at 16 fps. Panel (b) reports component latency for 2048×1152 outputs with 241 frames at 24 fps. Each stacked bar shows VAE encoding, denoising, and VAE decoding. The first bar is the three-step LTX-2.3 Refiner baseline. The second bar is the one-step SoL-Refiner, and the final two bars add TAE and Sol-Engine. Red arrows report the speedup between adjacent configurations. The final configuration reduces refinement latency from 57.461 to 6.447 seconds, an 8.91× speedup over the baseline. Lower latency is better.
Figure 7 : H3 acceleration on a single NVIDIA DGX Spark (GB10). (a) End-to-end latency for full-resolution four-step H3, its quantized variant, SoL-H3, and SoL-H3 with SoL-Refiner. Full-res denotes direct 768p generation; both SoL-H3 configurations generate at 384p and refine to 768p. Speedups use the 374-second full-resolution four-step baseline. (b) Stage-wise latency for the two SoL-H3 configurations. Stage 1 takes 21.60 seconds in both; Stage 2 decreases from 34.57 to 18.07 seconds with SoL-Refiner. The red annotation in panel (b) reports the 16.50-second reduction in total latency (29.4% relative to SoL-H3). The panels use separate linear scales.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Resolution
Configuration
GPU
VBench ↑
Latency (s) ↓
SC
BC
MS
DD
AQ
IQ
AVG
81 frames
161 frames
Base (Multi-Step)
H100
0.9082
0.9364
0.9813
0.7067
0.5618
0.6713
0.7943
127.7
490
2K
w/ Sol-Engine
H100
0.9092
0.9365
0.9813
0.7200
0.5550
0.6601
0.7937
53.7
164
Base (Multi-Step)
H100
0.9037
0.9367
0.9806
0.7267
0.5547
0.6719
0.7957
927.6
5123
4K
w/ Sol-Engine
H100
0.9043
0.9386
0.9817
0.7267
0.5511
0.6591
0.7936
284.8
1650
Appendix
Table 6 : Latency and quality comparison. Measurements use one H100 GPU. VBench metrics are evaluated on the corresponding 81-frame videos. Latency is reported in seconds per sample.
Source Generator
Videos
Native Res.
Aligned Crop
Main Refiner Input
WAN 2.1 1.3B
50
832×480
832×448
1024×576
SANA-Video 2B
50
1280×736
1280×704
1024×576
LTX-2 Stage 1
50
768×512
768×512
1024×576
Appendix
Table 7 : Source preparation for the main 2K Refiner-Bench setting. Native and aligned resolutions are generator-specific, while every aligned video is resized to the shared 1024×576 refiner input.
Content Group
Videos
Main Visual Challenges
Human daily activity
18
Hands, object interaction, water, and reflections
Animals
15
Fur, articulated motion, fast motion, and water
Nature and landscape
15
Water, weather, foliage, and particles
Sports and dance
15
Fast articulated motion, hands, and fine body geometry
Cinematic portraits
12
Faces, skin, hair, hands, and low-light detail
Urban and architecture
12
Structural lines, repeated geometry, and reflections
Appendix
Table 8 : Content groups in Refiner-Bench. The descriptions summarize the main visual challenges represented by each group.
Generator
Configuration
Resolution
VBench ↑
UniPercept ↑
Latency ↓
Overhead ↓
WAN-5B
Direct high-res
1280×704
80.15
54.18
87.70
–
Low-res base
832×480
79.43
48.14
33.59
–
Low-res + SoL-Refiner
1280×704
81.22
54.49
39.75
6.16
WAN-1.3B
Direct high-res
1280×720
79.73
55.08
272.69
–
Low-res base
832×480
80.40
52.42
72.55
–
Low-res + SoL-Refiner
1280×704
81.70
56.66
78.71
6.16
Appendix
Table 9 : Detailed quality and latency results across base generators. VBench is reported on a 0–100 scale. Latency and SoL-Refiner overhead are measured in seconds per sample on one H100 GPU. The overhead is the red Stage-2 segment in Figure 4 .
Generator
Configuration
SC ↑
BC ↑
MS ↑
DD ↑
AQ ↑
IQ ↑
AVG ↑
WAN-5B
Direct high-res
0.92258
0.95235
0.98199
0.71333
0.58522
0.65348
0.80149
Low-res base
0.90244
0.92967
0.96584
0.82000
0.53374
0.61434
0.79434
Low-res + SoL-Refiner
0.90250
0.94432
0.96426
0.82000
0.56545
0.67668
0.81220
WAN-1.3B
Direct high-res
0.92994
0.95195
0.97886
0.71333
0.57893
0.63057
0.79727
Low-res base
0.92382
0.94247
0.97691
0.78000
0.57314
0.62767
0.80400
Low-res + SoL-Refiner
0.92213
0.95378
0.97583
0.79333
0.59331
0.66346
0.81697
Appendix
Table 10 : Detailed VBench results across base generators. SC, BC, MS, DD, AQ, and IQ denote subject consistency, background consistency, motion smoothness, dynamic degree, aesthetic quality, and imaging quality. All metrics are reported on their original 0–1 scale.
Generator
Configuration
IAA ↑
IQA ↑
ISTA ↑
AVG ↑
WAN-5B
Direct high-res
60.6266
61.3029
40.5970
54.1755
Low-res base
52.4973
53.1629
38.7562
48.1388
Low-res + SoL-Refiner
59.1759
64.2542
40.0342
54.4881
WAN-1.3B
Direct high-res
62.5408
62.5327
40.1709
55.0815
Low-res base
59.1791
58.7869
39.2797
52.4152
Low-res + SoL-Refiner
63.1862
66.1531
40.6515
56.6636
Appendix
Table 11 : Detailed UniPercept results across base generators. IAA, IQA, and ISTA denote Image Aesthetics Assessment, Image Quality Assessment, and Image Structure and Texture Assessment, respectively.
Figure 8 : Additional qualitative effects of RL post-training. Panels (a) and (b) show selected human and animal examples. Each panel compares the low-quality input, the refiner without RL post-training, and the same refiner with RL post-training at four aligned frames from 0.0 to 5.0 seconds. Each complete frame is followed by a magnified crop from the red box. RL post-training sharpens the side profile and cap edges in panel (a), and improves the eyes, facial contours, and fur texture in panel (b).