PAIR: Bridging Perception and Action in Vision-Language-Action Models
Organizations: University of Maryland, College Park
Abstract
Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.
Figures & tables
| Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| SpatialVLA ( Qu et al., 2025 ) | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 |
| WorldVLA ( Cen et al., 2025 ) | 87.6 | 96.2 | 83.4 | 60.0 | 81.8 |
| -FAST ( Pertsch et al., 2025 ) | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 |
| LARA (full) ( Liu et al., 2026b ) | 88.0 | 92.0 | 88.5 | 86.0 | 88.6 |
| GR00T N1 ( Bjorck et al., 2025 ) | 94.4 | 97.6 | 93.0 | 90.6 | 93.9 |
| ( Black et al., 2025 ) | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 |
| Method | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
|---|---|---|---|---|---|---|---|---|
| OpenVLA ( Kim et al., 2025b ) | 0.8 | 3.5 | 23.0 | 8.1 | 34.8 | 15.2 | 28.5 | 15.6 |
| WorldVLA ( Cen et al., 2025 ) | 0.1 | 27.9 | 41.6 | 43.7 | 17.1 | 10.9 | 38.0 | 25.0 |
| NORA ( Hung et al., 2025 ) | 2.2 | 37.0 | 65.1 | 45.7 | 58.6 | 12.8 | 62.1 | 39.0 |
| UniVLA ( Bu et al., 2025 ) | 1.8 | 46.2 | 69.6 | 69.0 | 81.0 | 21.2 | 31.9 | 42.9 |
| ( Black et al., 2025 ) | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 |
| VLA-Adapter ( Wang et al., 2026 ) | 36.2 | 37.9 | 74.6 | 70.6 | 76.1 | 58.0 | 69.7 | 59.1 |
| Method | 1 | 2 | 3 | 4 | 5 | Avg. Len. |
|---|---|---|---|---|---|---|
| OpenVLA ( Kim et al., 2025b ) | 91.3 | 77.8 | 62.0 | 52.1 | 43.5 | 3.27 |
| RoboVLMs ( Li et al., 2026 ) | 98.0 | 93.6 | 85.4 | 77.8 | 70.4 | 4.25 |
| VPP ( Hu et al., 2025 ) | 96.5 | 90.9 | 86.6 | 82.0 | 76.9 | 4.33 |
| UnifiedVLA ( Wang et al., 2025b ) | 98.9 | 94.8 | 89.0 | 82.8 | 75.1 | 4.41 |
| DreamVLA ( Zhang et al., 2025 ) | 98.2 | 94.6 | 89.5 | 83.4 | 78.1 | 4.44 |
| NIAF ( Liu et al., 2026a ) | 99.7 | 95.9 | 90.6 | 84.8 | 76.4 | 4.47 |
| Position | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| Start (base) | 99.2 | 99.6 | 98.6 | 96.2 | 98.4 |
| End | 98.4 | 99.0 | 97.6 | 94.6 | 97.4 |
| Middle | 98.8 | 98.9 | 98.3 | 95.6 | 97.9 |
| All | 98.7 | 99.0 | 97.9 | 94.8 | 97.6 |
| Spatial | Object | Goal | Long | Avg. | |
|---|---|---|---|---|---|
| 0.1 (base) | 99.2 | 99.6 | 98.6 | 96.2 | 98.4 |
| 0.01 | 98.8 | 99.2 | 98.0 | 95.2 | 97.8 |
| 0 | 98.2 | 98.8 | 97.4 | 94.4 | 97.2 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Alignment Target | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| Raw expert actions | 98.2 | 98.6 | 97.0 | 94.4 | 97.1 |
| Action Latent Tokens (base) | 99.2 | 99.6 | 98.6 | 96.2 | 98.4 |
| Injection Strength | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| Learnable (base) | 99.2 | 99.6 | 98.6 | 96.2 | 98.4 |
| 1.0 | 98.8 | 99.4 | 98.2 | 95.6 | 98.0 |
| 0.5 | 98.9 | 99.3 | 98.1 | 95.9 | 98.1 |
| AE Training | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| Noise (base) | 99.2 | 99.6 | 98.6 | 96.2 | 98.4 |
| Clean | 98.4 | 99.0 | 97.6 | 95.4 | 97.6 |
| Stage / backbone | Batch | Steps | Learning rate | Schedule | LoRA rank |
|---|---|---|---|---|---|
| Action Autoencoder | 64 | 50k | cosine, 1k warmup | – | |
| VLA-Adapter + PAIR | 64 | 150k | cosine | 64 | |
| OpenVLA-OFT + PAIR | 64 | 50k | constant | 32 |
| Task | Instruction and success criterion | Demos. |
|---|---|---|
| Object placement | “Place the object in the bowl.” Place the object in the instructed bowl. | 51 |
| Button pressing | “First press , then , and finally .” Press the three buttons in the specified order. | 50 |
| Cube placement | “Put in the cup.” Select the instructed cube among the candidates and place it in the cup. | 51 |
| Cup stacking | “Stack on top of .” Stack the instructed source cup on the target cup. | 51 |
| Board wiping | “Wipe the marked circle off the whiteboard.” Grasp the eraser and remove the marked circle. | 30 |
| Water pouring | “Pour water from the red cup into the blue cup.” Transfer the water between the specified cups. | 30 |