Brain-Conditioned Action Policies for Neural Motor Decoding
Authors: Luyao Jin, Running Zhao, Huan Zhao, Vincent C. K. Cheung, Wei-Hsin Liao
Organizations: Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong · Department of Electrical and Computer Engineering, University of Hong Kong · School of Biomedical Sciences, The Chinese University of Hong Kong
Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose BrainVLA, a framework that enables neural motor decoding by drawing on a pretrained vision-language-action (VLA) model through language-mediated alignment. BrainVLA mitigates reliance on scarce paired neural-action data by leveraging VLA policies. We first construct VLA-compatible datasets including paired neural activity, action signals, language instructions, and rendered visual observations. Then, we adapt the OpenVLA-OFT policy to the target action spaces through LoRA fine-tuning. To establish an effective interface through which neural activity can convey motor intention to adapted VLA policies and guide action generation, we train a neural encoder via neural-language alignment, using language representations as semantic targets to capture latent motor intent from neural activity. The resulting neural representations serve as an endogenous intention signal to guide VLA policies to generate executable actions, while visual observations provide complementary information about the evolving task state. BrainVLA is evaluated on two neural motor datasets with different action dimensionalities using causal rollout decoding. It outperforms the evaluated baselines in cross-session decoding R2 and task success rate, while demonstrating high training data efficiency. These results establish a route for neural motor decoding to draw on large-scale robotic priors through brain-conditioned VLA policies.
Figures & tables
Figure 1: Overview of BrainVLA. (A) The construction of the VLA-compatible dataset, with MuJoCo-rendered observations, language instructions, and task-specific success criteria. (B) The task-specific VLA adaptation, with language, vision, and proprioception, while task-specific modules and LoRA parameters are optimized. (C) Train a neural encoder to map neural activity into the VLA conditioning space through pooled neural-language alignment.
Figure 2: BrainVLA causal rollout protocol. At each step, a causal neural window and the current MuJoCo-rendered observation are used to predict action A^t . Executing A^t updates the environment and produces the next rendered observation ot+Δ , which is fed back to BrainVLA while the neural window advances. Thus, inference uses only past neural activity and self-evolving visual observations generated by the model’s own actions.
Method
D1
D2
R2
SR
R2
SR
Wiener filter ( Glaser et al. (2020) )
0.15 ± 0.11
0.13 ± 0.08
0.14 ± 0.11
0.19 ± 0.06
RNN ( Glaser et al. (2020) )
0.32 ± 0.16
0.21 ± 0.17
0.18 ± 0.21
0.26 ± 0.12
CycleGAN+WF ( Ma et al. (2023) )
-0.09 ± 0.15
0.07 ± 0.05
-0.09 ± 0.10
0.12 ± 0.03
POYO ( Azabou et al. (2023) )
0.14 ± 0.22
0.34 ± 0.14
0.12 ± 0.18
0.24 ± 0.09
NDT2 ( Ye et al. (2023) )
0.21 ± 0.18
0.16 ± 0.13
0.23 ± 0.14
0.24 ± 0.12
Table 1: Decoding performance comparison. The best results are highlighted in bold.
Figure 3: Representative causal BrainVLA rollouts on D1 (left panel) and D2 (right panel). Each pane shows sequential observations from one neural-conditioned episode.
Variant
Decoder
Alignment
Vision
D1
D2
R2
SR
R2
SR
BrainVLA
adapted VLA
✓
Evolving
0.65
0.77
0.46
1.00
w/o VLA adapt.
OpenVLA-7B
✓
Evolving
0.39
0.43
0.34
0.75
w/o language
adapted VLA
✗
Evolving
0.56
0.68
0.41
0.87
w/o vision
adapted VLA
✓
White
0.29
0.52
0.13
0.44
w/o VLA
MLP
✓
Evolving
0.14
0.31
0.18
0.22
Table 2: Component ablation. Each variant is retrained and reevaluated to assess the contributions.
Figure 4: (A) R2 and task success rate with scaled training data. BrainVLA consistently outperforms POYO and LSTM across data regimes and remains effective under limited neural supervision. Shaded regions indicate variability across test sessions. (B) R2 for each of the seven action dimensions under the full model and w/o vision or neural input. Vision removal primarily degrades the translational dimensions (x,y,z) , whereas neural removal more strongly affects roll and finger-related dimensions.
Variant
Neural
Vision
Proprio
D1
D2
R2
SR
R2
SR
BrainVLA
✓
✓
✓
0.65
0.77
0.46
1.00
Neural+Vision
✓
✓
✗
0.35
0.39
0.45
0.99
Neural+Proprio
✓
✗
✓
0.27
0.29
-0.04
0.03
Vision+Proprio
✗
✓
✓
0.24
0.25
0.17
0.87
Neural
✓
✗
✗
0.15
0.11
-0.06
0.43
Table 3: Modality ablation. The trained BrainVLA model is kept fixed while removing modalities at inference to measure their individual and complementary contributions.
rotate the wrist [direction] to align the hand with the red item
Shape
move the thumb inward spread the thumb
Grasp
[action] the red item
Carry
carry the red item toward the green target while [holding] it
Orient2
keep [holding] the red item and rotate the wrist [direction] to align with the green target
Appendix
Table 5: D1 language instructions. [direction] is clockwise or counter-clockwise ; [action] is pinch , grasp , or scoop ; and [holding] is pinching , grasping , or scooping . All six holding-direction combinations occur under Orient2 .”
Phase
Group / pattern
Instruction template
Initialization
Both
hold both finger groups at their corresponding colored center target markers
Outward
Index
[motion] the index finger toward the red target marker while keeping the MRS finger group centered
Outward
MRS
[motion] the MRS finger group toward the yellow target marker while keeping the index finger centered
Outward
Both / same
[motion] both finger groups toward their corresponding colored target markers
Outward
Both / opposite
[motion] the index finger toward the red target marker and [opposite] the MRS finger group toward the yellow target marker
Return
Index
[motion] the index finger back toward the red center target marker while keeping the MRS finger group centered
Appendix
Table 6: D2 suggested language instructions. [motion] is flex or extend . For opposite-direction movements, [opposite] is extend when [motion] is flex , and vice versa. Each movement template represents two observed strings.
Parameter
Setting
Simulation timestep
0.02 s
Gravity
(0,0,0)
Render size
256×256
Hand root position
(0.52,0.28,0.93)
x translation range
[−0.55,0.55]
y translation range
[−0.35,0.35]
Appendix
Table 7: MuJoCo scene configuration used for D1 reconstruction.
Parameter
Setting
Simulation timestep
0.02 s
Gravity
(0,0,0)
Render size
256×256
Controlled DoF
Index, MRS finger group
Index joint range
[0,1.35] rad
MRS joint range
[0,1.35] rad
Appendix
Table 8: MuJoCo scene configuration used for D2 reconstruction.
Task
Success criterion
Reach
mintdtitem≤0.055
Orient
dTitem≤0.075 and eTroll≤0.35
Shape
∣qT−qTGT∣≤0.18 and direction agreement as defined below
Grasp
mintdtitem≤0.065 and cT≥0.60
Carry
dTtarget≤0.070 and cT≥0.60
Orient2
dTtarget≤0.070 and eTroll≤0.35
Appendix
Table 9: D1 success criteria using the evaluator’s default thresholds. All conditions within a row must hold.
Parameter
Setting
Base model
OpenVLA-7B
Fine-tuning method
LoRA + action head + proprioceptive projector
Optimization steps
50,000
Effective batch size
8
Gradient accumulation steps
1
Optimizer
AdamW
Appendix
Table 10: OpenVLA fine-tuning configuration for D1.
Parameter
Setting
Base model
OpenVLA-7B
Fine-tuning method
LoRA + action head + proprioceptive projector
Optimization steps
50,000
Effective batch size
8
Gradient accumulation steps
1
Optimizer
AdamW
Appendix
Table 11: OpenVLA fine-tuning configuration for D2.
Setting
Value
Model dimension
256
Latent queries
38
Encoder latents
38 ( num_encoder_latents =0 )
Cross-attention
1 head, head dim 64, RoPE on Q/K/V
Self-attention depth
4
Self-attention heads
8, head dim 64 (inner dim 512)
Appendix
Table 12: Neural encoder configuration.
Parameter
Setting
Embedding dimension d
4096
Batch size B
32
Neural pooling
Mean over latent slots ni=M1∑m=1MZi,mN
Language pooling
Masked mean over valid instruction tokens ti=∑ℓmi,ℓ∑ℓmi,ℓTi,ℓ
<s> In: What action should the robot take to ; standard frozen OpenVLA token embeddings.
Neural slots
11–48
38
One 4096-dimensional embedding per slot, produced from causal spiking activity by the neural bridge. Placeholder lookup embeddings are completely replaced.
Text suffix
49–53
5
?\n Out:␣ ; standard frozen OpenVLA token embeddings.
Policy prompt
1–53
10+38+5=53
Fixed across tasks; no task-language tokens are supplied to the policy.
Action targets
54–109
56
Eight future actions × seven action dimensions, represented using OpenVLA action tokens during training.
Stop token
110
1
End-of-sequence supervision used during action training.
Appendix
Table 14: Composition of the fixed neural prompt supplied to OpenVLA. The task instruction is not included in the policy input.
Policy
D1 ( R2 )
D1 (SR)
D2 ( R2 )
D2 (SR)
OpenVLA-7B + task heads
0.86 ± 0.03
0.85 ± 0.06
0.39 ± 0.02
0.99 ± 0.01
OpenVLA-7B + LoRA + task heads
0.89 ± 0.02
0.97 ± 0.04
0.49 ± 0.02
1.00 ± 0.00
Appendix
Table 15: VLA adaptation with language and vision (oracle policy). Task head specifies learnable action heads and proprioception projectors.
Figure 5: Language-conditioned closed-loop rollout protocol. At each step, the fixed language instruction and the current MuJoCo-rendered observation are used to predict action A^t . Executing A^t updates the environment and produces the next rendered observation ot+Δ , which is fed back to the VLA policy. Thus, inference uses only past fixed language instruction and self-evolving visual observations generated by the model’s own actions.
Figure 6: Representative closed-loop BrainVLA rollouts on D1. Each row shows sequential observations from one neural-conditioned episode, progressing from left to right. At every step, BrainVLA predicts an action from the causal neural history and current rendered observation; the action is executed in MuJoCo to generate the next observation. The examples illustrate progressive reaching, object interaction, and target-directed manipulation under autoregressive visual feedback.
Figure 7: Representative closed-loop BrainVLA rollouts on D2. Each row shows sequential observations from one neural-conditioned episode, progressing from left to right. At every step, BrainVLA predicts an action from the causal neural history and current rendered observation; the action is executed in MuJoCo to generate the next observation. The examples illustrate progressive reaching, object interaction, and target-directed manipulation under autoregressive visual feedback.
Figure 8: Task-level analysis on D1. Task success rates under the full BrainVLA model and ablations without vision or neural input. Reach and Carry are more sensitive to vision removal, while other tasks rely more strongly on neural conditioning. Error bars denote variability across test sessions.
Figure 9: Comparison of joint and staged training on D1. Left: overall and dimension-wise action R2 for the 7-DoF action space. Right: overall and task-specific success rates. Staged training provides stronger and more consistent performance overall, supporting the sequential adaptation strategy used in BrainVLA.
Figure 10: R2 and task success rate with scaled training D2 data. BrainVLA consistently outperforms POYO and LSTM across data regimes and remains effective under limited neural supervision. Shaded regions indicate variability across test sessions.
Figure 11: Alignment between neural representation and language representation. Warmer colors denote higher representational similarity.
Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale, yet the action decoder, which converts the backbone's hidden state into a continuous control signal, has barely changed and remains a single-point predictor across the majority of current VLAs. Whether implemented via autoregressive token bins, L1 regression, or flow-matching denoising, the resulting decoder treats the action space as unstructured, leaving the geometric proximity of neighboring actions unexploited during training. To advance this, we introduce ActionMap, a voxel heatmap action head that drops into an existing VLA in place of its native action decoder. For each new action, the head predicts a voxel heatmap over the action space, where each voxel directly stores the probability of the corresponding action. Across LIBERO simulation and real-world Franka manipulation, our heatmap head surpasses two architecturally distinct backbones at matched training steps (e.g., +8.2% over OpenVLA-OFT's L1 regression head on the LIBERO four-suite average), converges at comparable or faster rates on both backbones, and remains markedly more data-efficient at low training data. The cross-backbone consistency indicates that action representation is a real lever for VLA performance, distinct from further backbone or recipe scaling. Project Page: https://showlab.github.io/ActionMap/.
Pei Yang, Hai Ci, Yanzhe Chen +3
1Show Lab, National University of Singapore · 2NVIDIA
Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs-convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neural Action Codec (NAC), which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture. We observe that audio-specific mel-spectrogram objectives are ill-suited for kinematic signals; however, by replacing them with simple time-domain and non-mel spectral reconstruction losses, audio-codec-style models can autoencode actions with high fidelity without substantial architectural changes. NAC provides a compact, ordered token space via offset codebooks, enabling standard autoregressive policies to operate over short, structured sequences. Meanwhile, a Vocos-style decoder with an ISTFT head and adversarial discriminators recovers smooth, detailed trajectories. Across LIBERO-10, RoboMimic, and a suite of real-world manipulation tasks, NAC achieves lower reconstruction error and higher success rates than binning, FAST, and prior VQ-based tokenizers at comparable or better compression rates. These results demonstrate that repurposed neural audio codecs offer a strong, practical backbone for learned action tokenization in modern VLAs.
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models' internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5× when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.