Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.
Figures & tables
Fig. 1 : We propose TOAST, a novel framework that introduces stochastic action tokenization to enable efficient policy learning with supervision diversification while preserving the underlying action.
Fig. 2 : TOAST tokenization pipeline. Given a discrete action chunk after quantization (e.g., DCT) and flattening, TOAST uses a unigram language model to calculate probabilities over the n -best token sequences for the chunk and then samples a token sequence. Tokenization sampling is performed only during training.
Fig. 3 : Evaluation tasks. We evaluate TOAST on the LIBERO benchmark and four real-robot manipulation tasks.
Fig. 4 : Our real robot setup. The robot (Franka Research 3) is equipped with a wrist camera (RealSense D405) and a Robotiq 2F-85 Gripper. Two scene cameras (RealSense D435) are placed facing the robot.
Table
Grocery
Breakfast
Drawer
Bussing
Bagging
Setup
Stowing
200
200
100
60
TABLE I : Amount of real-robot task demonstrations used for training policies.
LIBERO
Table
Grocery
Breakfast
Drawer
Bussing
Bagging
Setup
Stowing
16.15
38.47
41.00
41.37
39.48
TABLE II : Average episode duration of the datasets (in seconds).
Tokenizer
Fraction of LIBERO training data
1/1
1/2
1/4
1/8
1/16
Binning [ 5 ]
89.5
80.0*
70.7*
56.4*
37.8*
FAST+ [ 10 ]
85.7*
59.4*
40.6*
30.1*
17.8*
FAST (rebuilt) [ 10 ]
92.4
83.8
71.3*
55.7*
34.4*
BEAST [ 26 ]
91.5
82.6
69.8*
55.2*
37.1*
VQ-VLA [ 13 ]
82.8*
74.7*
58.3*
46.8*
31.6*
TABLE III : LIBERO Results with different amounts of training data. Mean success rate is calculated based on 2,000 rollouts per seed. * denotes that TOAST is significantly better than this row (the 95% bootstrap CI of the difference excludes zero). This shows that stochastic action tokenization is more effective as available training data is reduced.
Fig. 5 : Real-robot task results. We evaluate policies with three different tokenizations: TOAST, TOAST (deterministic), and FAST+. Each policy is evaluated based on 30 rollouts per task with three random seeds. TOAST achieves the highest success rate (50.8%), improving performance of TOAST (deterministic) by 15.8 points and FAST+ by 20.0 points.
Fig. 6 : LIBERO ( 1/8 ) results with varying sampling hyperparameter α . Performance is stable for α∈[0.05,1.0] and degrades toward the deterministic baseline for larger α .
Tokenizer
Fraction of LIBERO training data
1/1
1/2
1/4
1/8
1/16
(Vocabulary: DROID + LIBERO)
TOAST (deterministic)
90.1*
81.8*
69.2*
53.0*
36.4*
TOAST
92.2
85.3
74.7
61.3
46.0
(Vocabulary: LIBERO)
TOAST (deterministic)
89.1*
80.3*
67.5*
53.8*
37.9*
TABLE IV : Effect of the vocabulary corpus. LIBERO success rate of TOAST and TOAST (deterministic) with vocabularies built from DROID+LIBERO and only LIBERO. The LIBERO portion is restricted to the same 1/k subset used for policy training. ∗ : TOAST is significantly better than TOAST (deterministic).
Fig. 7 : Success rate gain by sampling. * denotes a significant improvement over the deterministic baseline. In both flattening directions, stochastic action tokenization tends to achieve more improvements when limited data is available.
Fig. 8 : Sensitivity to equivalent tokenizations ( Δtok=Lalt−Lcan ) on 100 held-out LIBERO demonstrations. Stochastic tokenization reduces dependence on a single tokenization of an action chunk.
Hyperparameter
LIEBRO
Real Robot (FR3)
Vocabulary size
512
512
α (smoothing parameter)
0.1
0.5
n ( n -best search)
64
64
TABLE V : Hyperparameters for TOAST.
Hyperparameter
LIBERO
Real Robot (FR3)
Training steps
30,000
30,000
Warmup steps
1,000
1,000
Batch size
32
32
Learning rate
2.5×10−5
2.5×10−5
Horizon
10
20
TABLE VI : Hyperparameters used for training policies.