Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly iterative sampling. To address these limitations, we unify regression and flow matching under a shared objective and extend it to derive a quantile objective. This quantile objective guides the design of our Quantile Head, which predicts a median and positive gaps to form ordered marginal action quantiles in one forward pass. These quantiles support multiple sampling strategies without retraining and are jointly supervised to train the default median policy. Our local analysis of this joint supervision shows that, with calibrated nearby quantiles, fixed gaps, and matched correction speed, direct median updates have lower variance than under median-only supervision. Experiments show that this jointly supervised median policy achieves the highest average success rates among the compared methods on LIBERO, LIBERO-Plus, LIBERO-Pro, and two real-robot tasks, together with the shortest mean episode time among matched LIBERO baselines; code is available at https://github.com/xwangrs/Quantile-Head-for-VLA.
Figures & tables
Figure 1: Action head comparison. Regression predicts a point, iterative flow matching reuses the same head at updated action states and integration times, and our head predicts ordered marginal quantiles in one pass for direct median control or optional sampling.
Figure 2: Architecture of the Quantile Head. Our model combines a frozen VLM, learnable prompts, and a Quantile Head. The VLM encodes images and instructions. Learnable prompts precede action tokens; masked self-attention lets prompts read VLM features and action tokens read both VLM and prompt features. The action head maps the resulting action features to a median and positive gaps, then accumulates these gaps around the median to form ordered marginal quantiles. We supervise these quantiles with masked pinball loss. Inference uses the median by default or samples from the predicted marginal quantiles.
Figure 3: Attention and loss masks (schematic). (a) Block self-attention controls information flow among VLM, prompt, and action tokens. (b) Coordinate masks are independently resampled for each example in every training batch and shared across all quantiles.
Table 1: Success rates (%) on LIBERO and LIBERO-Pro. Best in bold ; second-best underlined . Superscripts indicate result sources: a Kim et al. (2025) ; b Wu et al. (2026b) .
(a) Zero-shot transfer
Method
Camera
Robot
Lang.
Light
Bkg.
Noise
Layout
Total
OpenVLA a ( Kim et al., 2024 )
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
π0 -FAST a ( Pertsch et al., 2025 )
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
π0 a ( Black et al., 2024 )
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
OpenVLA-OFT a ( Kim et al., 2025 )
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
π0.5 b ( Black et al., 2025 )
75.8
79.4
83.3
95.5
95.0
89.6
87.0
85.7
Table 2: Success rates (%) on LIBERO-Plus. Within each setting, best in bold ; second-best underlined . Superscript sources: a Fei et al. (2026) ; b Zhong et al. (2026) ; d Shi et al. (2026a) . Unmarked baselines use their original papers.
Real-robot Task
π0.5
Quantile Head
Δ
Place apple on yellow plate
31
52
+21
Remove cuboid from blue plate
86
98
+12
Average
58.5
75.0
+16.5
Table 3: Real-robot success rates (%). The π0.5 baseline is fully fine-tuned. Δ gives the gain over π0.5 in percentage points.
Setting
Spatial ↑
Object ↑
Goal ↑
Long ↑
Avg. ↑
(a) Supervised Fine-Tuning
Avg. Episode Time (s) ↓
L2
8.73
97.0
98.0
96.0
93.0
96.0
L1
8.13
99.0
99.0
96.0
95.0
97.3
Flow Matching
12.57
97.0
98.0
96.0
94.0
96.3
Quantile Head
5.88
99.0
100.0
99.6
98.6
99.3
(b) Prompt + Label Mask
Table 4: Ablations on the four LIBERO suites: success rates (%) and episode times (s). Episode times in (a) average over all evaluation episodes. See Appendix for ablation settings.
Figure 4: Inference-time sampling on LIBERO-Long. Colors denote quantile windows; bars show success-rate changes from median decoding (98.6%; dashed lines) in percentage points. Quantile (a) and Shared per chunk (b) use the same setting ( ρ=0 ). See Appendix .
Figure 5: Mean predicted action spread during manipulation. Median-centered Q0 – Q1 ranges averaged over six normalized action dimensions ( h=0 ); endpoints use linear tail extrapolation.
Figure 6: Training demonstrations and action head outputs. Top views show demonstration states; points show initial (dx,dy) commands from one training run, colored by outcome. Subcaptions report success rates; ×100 marks overlapping points. Details: Appendices – .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Number of quantiles on LIBERO-Long. Success rate as the number of predicted quantiles K varies. Rates average three training seeds, with 1,000 evaluation episodes per seed and setting. The default K=21 attains the highest observed rate of 98.6%.
Figure 8: Label mask ratio on LIBERO-Long. Success rate as the fraction of masked action labels r increases. Rates average three training seeds, with 1,000 evaluation episodes per seed and setting. The sweep includes the fully masked setting r=1.0 .
Suite
Original
Object
Position
Semantic
Task
Avg.
Spatial
99.0
100.0
49.0
98.0
47.0
73.5
Object
100.0
95.0
26.0
99.0
10.0
57.5
Goal
95.0
81.0
34.0
97.0
27.0
59.8
Long
98.0
67.0
16.0
97.0
20.0
50.0
Mean
98.0
85.8
31.3
97.8
26.0
60.2
Appendix
Table 5: Quantile Head success rates (%) by suite and perturbation type on LIBERO-Pro.
(a) Zero-shot transfer
Suite
Camera
Robot
Lang.
Light
Bkg.
Noise
Layout
Total
Spatial
75.8
89.4
97.2
100.0
99.6
98.6
98.4
93.7
Object
87.1
73.6
93.5
100.0
100.0
99.1
92.3
91.5
Goal
78.2
79.0
68.5
91.8
95.4
94.7
73.2
81.7
Long
46.8
81.9
92.7
92.0
94.5
87.8
88.8
82.1
Overall
71.6
80.7
87.6
96.1
97.2
94.8
87.8
87.1
Appendix
Table 6: Quantile Head success rates (%) on LIBERO-Plus by suite and perturbation. Overall uses instance-weighted aggregation across the four suites.
Figure 9: Predicted action spread across six motion dimensions. Median-centered ranges in environment-command units from the rollout in Figure ; rotation panels use fivefold vertical magnification.
Figure 10: Real-robot platform. The experimental setup and arm assembly, with the AgileX Piper arm, two-finger gripper, wrist RGB camera, and fixed external RGB camera labeled.
Figure 11: Two real-robot tasks. Frames from the task videos show apple placement and cuboid removal. Numbered keyframes indicate temporal order, with source-video times in seconds. Blue circles mark the manipulated objects, and arrows schematically indicate the direction of motion.
Figure 12: Two training demonstrations with the same initial observation. Blue and orange denote left and right detours under the shared instruction “pick up the alphabet soup and place it in the basket.” (a) Measured end-effector paths for t=0,…,65 , resolved into rightward and forward displacement relative to the initial approach direction. The gray circle marks the shared start, the black star the initial target position, and colored dots the largest simultaneous XY separation (22.45 cm at t=29 ). (b) Raw controller commands: stars mark the first actions and hollow diamonds the ten-step means. (d), (h) Top-down renderings of the recorded t=29 states, with screen right aligned to the rightward axis in (a). Dashed lines mark the shared start–target centerline; colored rings mark the projected end effector and arrows show its lateral offset in the restored state. All other frames are original training-camera observations. Here t counts executed actions. Both 315-step demonstrations end in task success.
Figure 13: Action distributions learned across six motion dimensions. Rows compare the four heads and columns show first-step action commands from the no-obstacle setting in Figure . Every panel uses 1,000 samples from one illustrated training run; bins and axes are shared across heads within each column. The vertical axis is probability mass per bin (%); stems denote point masses. Blue and orange dashed lines mark the two demonstrations, with a gray line where their values coincide. Only dx and dy have distinct demonstration targets at this step; the other four targets are zero. A narrow sampled distribution can occupy a single bin and need not be a point mass.
Setting
Success rate (%)
Quantile: native ranks
94.0
Quantile: independent ranks
87.0
Quantile: shared ranks
92.0
Quantile: median
100.0
Flow Matching (10 integration steps)
3.0
L1
100.0
Appendix
Table 7: Closed-loop success rates in the two-demonstration experiment. Evaluation uses the same training initial state. These rates measure completion at this fixed state.
Figure 14: Initial action commands and closed-loop outcomes. Points show one training run’s 100 episodes per decoder; subcaptions report success rates. Each point represents one episode’s first executed action. Green circles indicate eventual success; red crosses indicate failure. All panels use the same axes and, within this experiment, the same initial observation. Axis limits also match Figure for comparison with the obstacle setting. The 100 points in each of (e), (f), and (g) coincide exactly; no coordinate jitter is added. Colors label whole episodes, rather than the correctness of individual action coordinates.
Figure 15: Two successful training demonstrations around a central obstacle. Blue and orange denote left and right detours from the same initial observation under the shared instruction “pick up the alphabet soup and place it in the basket.” (a) Measured end-effector approach paths for t=0,…,105 ; the shaded rectangle is the footprint of the 6×4.4×19 cm obstacle. The gray circle marks the shared start, the black star the initial target position, and colored dots the largest simultaneous XY separation (30.24 cm at t=40 ). (b) Raw controller commands: stars mark the first actions and hollow diamonds the ten-step means. (d), (h) True overhead renderings of the saved t=40 states, with screen right aligned to the rightward axis in (a). Dashed centerlines and colored end-effector rings and offset arrows distinguish the two routes; the red dashed outline marks the obstacle’s projected top face, including occluded edges. All other frames are original training-camera observations. Here t counts executed actions. Both 430-action demonstrations finish successfully without recorded obstacle contact and are the only training trajectories for this experiment.
Figure 16: Conditional action distributions with a central obstacle. Rows compare four heads; columns show six motion coordinates of the first predicted action, using 1,000 samples per head from one illustrated training run at a shared initial observation. Each column shares 32 bins and axes. The vertical axis gives probability mass per bin (%) on a 0–105% scale. Stems denote exactly constant outputs; narrow distributions remain histograms. Blue and orange dashed lines mark the two demonstrations, with gray where references coincide: dz=−0.6 and all rotation targets zero. Quantile Head uses one shared rank τ=Φ(z) across motion coordinates; Flow Matching uses ten integration steps. These marginal distributions do not establish learned joint dependence or measure closed-loop success.
Success rate (%)
Setting
Earlier ( )
Obstacle
Quantile: native ranks
94.0
1.0
Quantile: independent ranks
87.0
4.0
Quantile: shared ranks
92.0
6.0
Quantile: median
100.0
0.0
Flow Matching (10 steps)
3.0
0.0
Appendix
Table 8: Task success rates in the original and obstacle settings. Each head is trained for 6,000 updates on its setting’s two demonstrations. These rates measure completion at a fixed training initial state.
Figure 17: Initial action commands and outcomes with a central obstacle. Points illustrate one training run; subcaptions report success rates. Each point is one episode’s first action: green circles indicate eventual success and red crosses indicate failure. All settings share the same obstacle-task initial observation. Axes match Figure for visual comparison; coordinates are environment action commands, not physical displacements. The 100 L1 and L2 points and the 100 median points each coincide exactly within their respective panels. No coordinate jitter is added. Colors describe whole episodes, rather than the correctness of an individual action.
Figure 18: Shared quantile sampling intervals and obstacle-task outcomes. Each panel illustrates 100 episodes from one training run at the same initial observation; subcaptions report success rates. Each point is the first executed action command: green circles indicate eventual episode success, red crosses indicate failure, and blue diamonds mark the two training demonstrations’ first actions. The six motion coordinates share one rank; the gripper uses its median. All panels use identical axes in environment action-command units, with no jitter or removal of coincident points. The deterministic median repeats are not independent task instances.
Selection
Scored coordinates
Gripper rank
Success (%)
Random, squared density
Six motion
Median
13
Random, linear density
All seven
Shared sampled rank
14
Maximum mean density
All seven
Shared selected rank
0
Appendix
Table 9: Density-based decoding in the obstacle setting. All variants use the same trained Quantile models.
Decoder
No obstacle
Obstacle
Quantile: uniform shared ranks
92
6
Quantile: density-weighted shared ranks
95
13
Quantile: median
100
0
Flow Matching (10 integration steps)
3
0
L1
100
0
L2
100
0
Appendix
Table 10: Two-demonstration task success rates (%). Evaluation uses each setting’s fixed training initial state. Uniform shared ranks cover [0,1] ; density weighting uses six motion coordinates with exponent two. Deterministic repeats are not independent task instances.