Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of π0 with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.
Figures & tables
Fig. 1: Overview of FRAM. (1) Observation inputs, (2) language and vision encoders, (3) prediction of the future end-effector trajectory, (4) trajectory-guided local feature extraction, and (5) action generation by flow matching.
Model
Size
Spatial
Object
Goal
Long
Ave.
OpenVLA-OFT [ 10 ]
7.71B
97.6
98.4
97.9
94.5
97.1
FLOWER [ 13 ]
950M
97.5
99.1
96.1
94.9
96.9
π0 [ 2 ]
3.3B
96.8
98.8
95.8
85.2
94.2
FRAM (Our)
138M
92.2
96.2
93.8
86.4
92.2
SmolVLA [ 12 ]
450M
90.0
96.0
92.0
71.0
87.3
π0 -FAST [ 22 ]
3B
96.4
96.8
88.6
60.2
85.5
TABLE I: Success rates (%) on standard LIBERO
Model
Cam.
Robot
Lang.
Light
Bk.
Noise
Layout
Ave.
OpenVLA-OFT [ 10 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
FRAM (Ours)
68.6
40.2
70.7
87.1
68.5
72.4
68.3
67.3
π0 -FAST [ 22 ]
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
π0 [ 2 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
OpenVLA [ 1 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
FRAM -Feat.
61.6
33.1
55.0
81.4
67.9
57.1
56.4
57.8
TABLE II: Success rates (%) per perturbation on LIBERO-Plus
Fig. 2: Examples of FRAM under LIBERO-Plus perturbations. (a) Background texture, (b) camera viewpoint, (c) object layout, and (d) robot initial pose are changed. The dots show the predicted trajectory, the end dot shows the target location, and the heatmap shows the local features.
Fig. 3: Examples from the same initial state. (a) FRAM, (b) FRAM − Feat., and (c) FRAM − Traj.. The top row is the external view and the bottom row is the wrist view. The dots show the predicted keypoint trajectory and the heatmap shows the visual features used for action generation. (b) predicts the trajectory but uses no image features, so no heatmap is shown. (c) uses the global ResNet feature map, shown as a heatmap over the whole image, but predicts no trajectory.
Fig. 4: Cup stacking with a dual-arm UR5e. (a) Left arm only, (b) right arm only, and (c) both arms (c1: left wrist, c2: right wrist). The white line is the predicted trajectory, the star is its end point, and the yellow frame marks the arm in motion.
Robot manipulation uses temporal context to select actions and visual foresight to assess their consequences, yet dense representations of past and future observations incur substantial processing costs. We introduce PACT-WAM, a world-action model that jointly generates a 16-step action trajectory and its temporally corresponding visual forecast through conditional flow sampling. Hierarchical history encoding assigns coarse spatial representations to earlier observations and finer representations to recent ones, retaining 16 observations with 256 tokens per view, 75% fewer than dense encoding of the same frames. A shared flow module jointly updates continuous action and visual states through two modality-specific heads under transition-wise causal attention, and a TiTok-VAE decoder reconstructs multi-view future images from the visual latents. Decoded forecasts also support Proposal Review (PR), a vision-language model component for execution-prefix selection and proposal rejection. Without PR, PACT-WAM achieves average success rates of 98.6%, 92.3%, and 78.0% on LIBERO, RoboTwin 2.0, and real-world Piper tasks, respectively. PR provides a test-time enhancement, raising these rates to 99.5%, 93.4%, and 86.7%. Ablations show that hierarchical history allocation and joint action-visual generation improve control success, while analyses of visual capacity and forecast-guided execution characterize the trade-offs between success and proposal-generation cost.
Yushan Liu, Jingjing Fan, Shoujie Li +3
Tsinghua University · Nanyang Technological University · Xspark AI
Mobile manipulation requires a robot to coordinate base and arm motion under continuously changing viewpoints and contact conditions, within an action space far larger than that of fixed-base manipulation. Existing Vision-Language-Action (VLA) policies are limited in two respects. (i)They map observations directly to whole-body action chunks, searching this large action space without an explicit task-space motion plan, which makes coordinated base--arm prediction imprecise. (ii)They execute the predicted chunk open-loop, without checking whether the actions can realize the motion the policy intended, so control errors and unmodeled contacts accumulate into a gap between planned and realized motion. We present DreamTrajectory, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation. Addressing(i), DreamTrajectory jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert, so that the trajectory explicitly guides base--arm action generation instead of remaining implicit. Addressing(ii), a lightweight trajectory world model predicts the trajectory that a candidate action chunk would induce, and a test-time search--predict--score procedure selects the candidate best aligned with the planned trajectory. On MS-HAB, trajectory guidance raises average success from 32.3% to 47.5% and test-time refinement further to 54.8%, with the largest gains on contact-rich articulated-object tasks. On three real-world mobile manipulation tasks, the corresponding average success rates are 63.3%, 81.7%, and 90.0%.
Zheng Yang, Wenjie Zhang, Xiangyu Chen +7
Hong Kong University of Science and Technology (Guangzhou) · Ola Dimensions
Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.
Mengya Liu, Baoxiong Jia, Jiangyong Huang +2
State Key Laboratory of General Artificial Intelligence, BIGAI · Peking University · Delta Intelligence