Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.
Figures & tables
Fig. 1 : We propose TOAST, a novel framework that introduces stochastic action tokenization to enable efficient policy learning with supervision diversification while preserving the underlying action.
Fig. 2 : TOAST tokenization pipeline. Given a discrete action chunk after quantization (e.g., DCT) and flattening, TOAST uses a unigram language model to calculate probabilities over the n -best token sequences for the chunk and then samples a token sequence. Tokenization sampling is performed only during training.
Fig. 3 : Evaluation tasks. We evaluate TOAST on the LIBERO benchmark and four real-robot manipulation tasks.
Fig. 4 : Our real robot setup. The robot (Franka Research 3) is equipped with a wrist camera (RealSense D405) and a Robotiq 2F-85 Gripper. Two scene cameras (RealSense D435) are placed facing the robot.
Table
Grocery
Breakfast
Drawer
Bussing
Bagging
Setup
Stowing
200
200
100
60
TABLE I : Amount of real-robot task demonstrations used for training policies.
LIBERO
Table
Grocery
Breakfast
Drawer
Bussing
Bagging
Setup
Stowing
16.15
38.47
41.00
41.37
39.48
TABLE II : Average episode duration of the datasets (in seconds).
Tokenizer
Fraction of LIBERO training data
1/1
1/2
1/4
1/8
1/16
Binning [ 5 ]
89.5
80.0*
70.7*
56.4*
37.8*
FAST+ [ 10 ]
85.7*
59.4*
40.6*
30.1*
17.8*
FAST (rebuilt) [ 10 ]
92.4
83.8
71.3*
55.7*
34.4*
BEAST [ 26 ]
91.5
82.6
69.8*
55.2*
37.1*
VQ-VLA [ 13 ]
82.8*
74.7*
58.3*
46.8*
31.6*
TABLE III : LIBERO Results with different amounts of training data. Mean success rate is calculated based on 2,000 rollouts per seed. * denotes that TOAST is significantly better than this row (the 95% bootstrap CI of the difference excludes zero). This shows that stochastic action tokenization is more effective as available training data is reduced.
Fig. 5 : Real-robot task results. We evaluate policies with three different tokenizations: TOAST, TOAST (deterministic), and FAST+. Each policy is evaluated based on 30 rollouts per task with three random seeds. TOAST achieves the highest success rate (50.8%), improving performance of TOAST (deterministic) by 15.8 points and FAST+ by 20.0 points.
Fig. 6 : LIBERO ( 1/8 ) results with varying sampling hyperparameter α . Performance is stable for α∈[0.05,1.0] and degrades toward the deterministic baseline for larger α .
Tokenizer
Fraction of LIBERO training data
1/1
1/2
1/4
1/8
1/16
(Vocabulary: DROID + LIBERO)
TOAST (deterministic)
90.1*
81.8*
69.2*
53.0*
36.4*
TOAST
92.2
85.3
74.7
61.3
46.0
(Vocabulary: LIBERO)
TOAST (deterministic)
89.1*
80.3*
67.5*
53.8*
37.9*
TABLE IV : Effect of the vocabulary corpus. LIBERO success rate of TOAST and TOAST (deterministic) with vocabularies built from DROID+LIBERO and only LIBERO. The LIBERO portion is restricted to the same 1/k subset used for policy training. ∗ : TOAST is significantly better than TOAST (deterministic).
Fig. 7 : Success rate gain by sampling. * denotes a significant improvement over the deterministic baseline. In both flattening directions, stochastic action tokenization tends to achieve more improvements when limited data is available.
Fig. 8 : Sensitivity to equivalent tokenizations ( Δtok=Lalt−Lcan ) on 100 held-out LIBERO demonstrations. Stochastic tokenization reduces dependence on a single tokenization of an action chunk.
Hyperparameter
LIEBRO
Real Robot (FR3)
Vocabulary size
512
512
α (smoothing parameter)
0.1
0.5
n ( n -best search)
64
64
TABLE V : Hyperparameters for TOAST.
Hyperparameter
LIBERO
Real Robot (FR3)
Training steps
30,000
30,000
Warmup steps
1,000
1,000
Batch size
32
32
Learning rate
2.5×10−5
2.5×10−5
Horizon
10
20
TABLE VI : Hyperparameters used for training policies.
Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long token sequences or learned latent tokenizers that lack structure, limiting their compatibility with downstream policies. In this work, we identify three desiderata for action tokenization - high compression, total decodability, and an ordered token space - and introduce Ordered Action Tokenization (OAT), a learned action tokenizer that satisfies all three. OAT discretizes action chunks into an ordered sequence of tokens using a transformer with registers, finite scalar quantization, and ordering-inducing training mechanisms. By training each token prefix to decode into a valid action chunk, OAT places coarse control information in early tokens and uses later tokens to refine residual detail, yielding an anytime tradeoff between inference cost and action fidelity. We validate OAT in two prevailing uses of action tokens: autoregressive policies that generate tokens for control, and token co-training policies that use token losses to shape the vision-language model context consumed by a flow-based action expert. Across three policy backbones and more than 60 tasks spanning five simulation benchmarks and real-world settings, OAT consistently delivers strong policy performance while offering significantly greater flexibility at inference time.
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Shijie Lian, Bin Yu, Zhaolong Shen +5
Huazhong University of Science and Technology · Zhongguancun Academy · DeepCybo +5
Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging. Existing tokenizers typically map each discrete code to a fixed continuous action prototype, ignoring the robot's current proprioceptive state. This limitation is particularly pronounced in manipulation, where the same action token may require different continuous controls under different joint configurations, object poses, and contact conditions. We therefore propose SA-VLA, a state-aware action tokenizer that conditions action decoding on robot state. We study two state-injection mechanisms for VQ-based action tokenization: cross-attention between state and action features, and a lightweight state adapter that predicts action-wise modulation factors for state-conditioned action modulation and reconstruction. The adapter formulation expands the effective support of a finite codebook by allowing each discrete token to represent a family of state-dependent continuous actions, while preserving the efficiency and compatibility of discrete action modeling. Integrated into an LLM-based VLA policy, SA-VLA supports both autoregressive and parallel action-token decoding with minimal changes to the model interface. On 12 RoboTwin manipulation tasks, SA-VLA improves the average success rate from 0.29 to 0.56 over the strongest tokenizer baseline. In zero-shot sim-to-real experiments on three real-world tasks, it further improves average success from 0.15 to 0.33 over the strongest tokenizer baseline. These results demonstrate that state-conditioned action decoding is a simple and effective mechanism for reducing the compression gap in discrete VLA policies.
Tengyue Jiang, Chunpu Xu, Jiayue Kang +1
2East China University of Science and Technology · 3Hong Kong Polytechnic University · 4Xi’an University of Electronic Science and Technology +1