Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts
Authors: Jingbo Wang, Wenxuan Song, Wenhao Yu, Han Zhao, Xi Wang, Jiayi Chen, Donglin Wang, Yan Wang, +1 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · South China University of Technology · University of Science and Technology of China · Westlake University · Zhejiang University · Tsinghua University
Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising steps. We introduce Discrete Forcing, a flow-matching framework that combines these representations through an explicit coarse-to-fine generation process. It first predicts discrete action tokens to establish a coarse action structure, then uses them to guide continuous action refinement. The discrete and continuous components share a common diffusion transformer backbone with specialized branches, maintaining a parameter count comparable to a conventional single-branch model while requiring only one forward pass per branch. Extensive evaluations across multiple benchmarks demonstrate improved performance and faster inference over a parameter-matched continuous action expert, with consistent performance gains as model capacity increases. Real-world experiments further demonstrate improvements on high-precision and dynamic manipulation tasks.
Figures & tables
Figure 1 : Overview of the Discrete Forcing (DF). (a) Discrete Forcing Expert (DFE) performs discrete action prediction followed by continuous refinement within a partially shared Action DiT conditioned on VLM features. (b) This two-stage design reduces inference cost while consistently improving performance across simulation and real-world benchmarks.
Figure 2 : Comparison of action generation paradigms. Existing methods rely on K -step iterative denoising in either (a) discrete or (b) continuous action spaces, whereas (c) Discrete Forcing (Ours) first predicts a coarse discrete action and then performs a single continuous refinement.
Method
Parameters
NFE ↓
Spatial
Object
Goal
Long
Avg.
Iterative Generation (NFE >2 )
DreamVLA ( Zhang et al., 2026 )
1.6B
10
97.5
94.0
89.5
89.5
92.6
MemoryVLA ( Shi et al., 2026 )
7B
10
98.4
98.4
96.4
93.4
96.5
π0.5 ( Physical Intelligence et al., 2025 )
3B
10
98.8
98.2
98.0
92.4
96.9
FlowerVLA ( Reuss et al., 2025 )
1B
8
97.5
99.1
96.1
94.9
96.9
Fast Generation (NFE ≤2 )
Table 1 : Comparison on the LIBERO benchmark. Prior methods are grouped according to their number of function evaluations (NFE): iterative generation with more than two NFE, and fast generation with NFE ≤2 . Best results are highlighted in bold.
Table 2 : Mean task success rate (%) on RoboTwin 2.0 and RoboCasa-GR1.
Method
Success Rate (%) ↑
NFE ↓
Action Latency (ms) ↓
End-to-End Latency (ms) ↓
P50
P95
P50
P95
StarVLA (NFE=10)
96.7
10
145.46
149.36
200.43
206.06
StarVLA (NFE=4)
96.1
4
58.05
59.88
119.86
123.57
StarVLA (NFE=2)
95.1
2
29.20
30.44
93.29
96.26
Discrete Forcing (Ours)
97.6
2
24.91
29.69
78.95
83.89
Table 3 : Inference efficiency comparison on LIBERO. Action latency measures the action-generation module only, while End-to-End Latency additionally includes the VLM forward pass. P50 and P95 denote the median and 95th-percentile latency, respectively.
Figure 3 : Model scaling experiments. Marker size denotes model parameter count, while the value in parentheses indicates the parameter count of the action expert. Our Discrete Forcing consistently improves with model scale while maintaining lower inference latency than the baseline.
Figure 4 : Comparison of the action generation mechanisms between our method and the continuous-generation baseline. DF denotes Discrete Forcing. DF@1 denotes the output after the first discrete forward pass, whereas DF@2 denotes the output after the subsequent continuous refinement. StarVLA@ n denotes the action output after the n -th iterative forward pass.
Figure 5 : Real-world experiments across four tasks.
Variant
Success Rate (%)
Δ
Discrete Forcing (Ours)
95.2
–
Generation Order
Discrete→Discrete
93.6
−1.6
Continuous→Continuous
91.8
−3.4
Architecture and Training Design
w/o Discrete-guided Source
92.6
−2.6
Table 4 : Ablation studies on LIBERO-Long. We analyze the effect of generation order, architectural designs, and discrete conditioning strategies.
Configuration
Setting
Model parameters
1.55B total
Action DiT
24 layers per execution path, with 2 shared and 22 branch-specific layers; hidden dim. 1024
Action horizon
8
Optimizer
AdamW, β=(0.9,0.95) , ϵ=10−8 , weight decay 10−8
Learning rate
1×10−5 for both VLM and Action Expert
Batch size
16 per GPU on 4 H100 GPUs; global batch size 64
Table 5 : Implementation details on LIBERO.
Task
Success Rate
Task
Success Rate
adjust_bottle
100
place_can_basket
56
beat_block_hammer
78
place_cans_plasticbox
32
blocks_ranking_rgb
26
place_container_plate
98
blocks_ranking_size
20
place_dual_shoes
60
click_alarmclock
64
place_empty_cup
96
click_bell
60
place_fan
34
Table 6: Task-level success rates (%) on RoboTwin 2.0. We report the success rate for each task, together with the average performance across all tasks.
Task
DiT4DiT
GR00T-N1.6
StarVLA- π
LangForce
DualCoT-VLA
LDA
Ours
BottleToCabinetClose
48.0
51.5
26.0
72.0
66.0
76.0
74.0
CanToDrawerClose
74.0
13.0
62.0
78.0
64.0
71.0
82.0
CupToDrawerClose
52.0
8.5
42.0
46.0
46.0
41.0
20.0
MilkToMicrowaveClose
50.0
14.0
50.0
56.0
58.0
52.0
62.0
PotatoToMicrowaveClose
36.0
41.5
42.0
36.0
30.0
41.0
40.0
WineToCabinetClose
42.0
16.5
32.0
46.0
38.0
57.0
56.0
Table 7 : Per-task success rate (%) on the RoboCasa-GR1 benchmark.
# Shared Layers
Success Rate (%)
0
88.2
2
95.2
4
95.0
6
93.6
8
93.4
10
90.0
Table 8 : Ablation on the number of shared Action DiT layers on LIBERO-Long.
Discrete Action ( α )
Gaussian Noise ( 1−α )
Success Rate (%)
0.3
0.7
95.2
0.5
0.5
93.2
0.7
0.3
91.0
Table 9 : Ablation on the composition of the discrete-guided continuous source on LIBERO-Long. α controls the contribution of the dequantized discrete action, while 1−α controls the Gaussian noise.