Brain-Conditioned Action Policies for Neural Motor Decoding
Organizations: Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong · Department of Electrical and Computer Engineering, University of Hong Kong · School of Biomedical Sciences, The Chinese University of Hong Kong
Abstract
Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose BrainVLA, a framework that enables neural motor decoding by drawing on a pretrained vision-language-action (VLA) model through language-mediated alignment. BrainVLA mitigates reliance on scarce paired neural-action data by leveraging VLA policies. We first construct VLA-compatible datasets including paired neural activity, action signals, language instructions, and rendered visual observations. Then, we adapt the OpenVLA-OFT policy to the target action spaces through LoRA fine-tuning. To establish an effective interface through which neural activity can convey motor intention to adapted VLA policies and guide action generation, we train a neural encoder via neural-language alignment, using language representations as semantic targets to capture latent motor intent from neural activity. The resulting neural representations serve as an endogenous intention signal to guide VLA policies to generate executable actions, while visual observations provide complementary information about the evolving task state. BrainVLA is evaluated on two neural motor datasets with different action dimensionalities using causal rollout decoding. It outperforms the evaluated baselines in cross-session decoding and task success rate, while demonstrating high training data efficiency. These results establish a route for neural motor decoding to draw on large-scale robotic priors through brain-conditioned VLA policies.
Figures & tables
| Method | D1 | D2 | ||
|---|---|---|---|---|
| SR | SR | |||
| Wiener filter ( Glaser et al. (2020) ) | 0.15 0.11 | 0.13 0.08 | 0.14 0.11 | 0.19 0.06 |
| RNN ( Glaser et al. (2020) ) | 0.32 0.16 | 0.21 0.17 | 0.18 0.21 | 0.26 0.12 |
| CycleGAN+WF ( Ma et al. (2023) ) | -0.09 0.15 | 0.07 0.05 | -0.09 0.10 | 0.12 0.03 |
| POYO ( Azabou et al. (2023) ) | 0.14 0.22 | 0.34 0.14 | 0.12 0.18 | 0.24 0.09 |
| NDT2 ( Ye et al. (2023) ) | 0.21 0.18 | 0.16 0.13 | 0.23 0.14 | 0.24 0.12 |
| Variant | Decoder | Alignment | Vision | D1 | D2 | ||
|---|---|---|---|---|---|---|---|
| SR | SR | ||||||
| BrainVLA | adapted VLA | Evolving | 0.65 | 0.77 | 0.46 | 1.00 | |
| w/o VLA adapt. | OpenVLA-7B | Evolving | 0.39 | 0.43 | 0.34 | 0.75 | |
| w/o language | adapted VLA | ✗ | Evolving | 0.56 | 0.68 | 0.41 | 0.87 |
| w/o vision | adapted VLA | ✓ | White | 0.29 | 0.52 | 0.13 | 0.44 |
| w/o VLA | MLP | ✓ | Evolving | 0.14 | 0.31 | 0.18 | 0.22 |
| Variant | Neural | Vision | Proprio | D1 | D2 | ||
|---|---|---|---|---|---|---|---|
| SR | SR | ||||||
| BrainVLA | 0.65 | 0.77 | 0.46 | 1.00 | |||
| Neural+Vision | ✗ | 0.35 | 0.39 | 0.45 | 0.99 | ||
| Neural+Proprio | ✗ | 0.27 | 0.29 | -0.04 | 0.03 | ||
| Vision+Proprio | ✗ | 0.24 | 0.25 | 0.17 | 0.87 | ||
| Neural | ✗ | ✗ | 0.15 | 0.11 | -0.06 | 0.43 | |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| D1 |
|---|
| held-in-calib_ses-19250108T111022 |
| held-in-calib_ses-19250119T114045 |
| held-in-calib_ses-19250120T115537 |
| D2 |
| held-in-calib_ses-2020-10-27-Run2_behavior+ecephys |
| Task label | Instruction template |
|---|---|
| Reach | reach toward the red item |
| Orient | rotate the wrist [direction] to align the hand with the red item |
| Shape | move the thumb inward spread the thumb |
| Grasp | [action] the red item |
| Carry | carry the red item toward the green target while [holding] it |
| Orient2 | keep [holding] the red item and rotate the wrist [direction] to align with the green target |
| Phase | Group / pattern | Instruction template |
| Initialization | Both | hold both finger groups at their corresponding colored center target markers |
| Outward | Index | [motion] the index finger toward the red target marker while keeping the MRS finger group centered |
| Outward | MRS | [motion] the MRS finger group toward the yellow target marker while keeping the index finger centered |
| Outward | Both / same | [motion] both finger groups toward their corresponding colored target markers |
| Outward | Both / opposite | [motion] the index finger toward the red target marker and [opposite] the MRS finger group toward the yellow target marker |
| Return | Index | [motion] the index finger back toward the red center target marker while keeping the MRS finger group centered |
| Parameter | Setting |
|---|---|
| Simulation timestep | s |
| Gravity | |
| Render size | |
| Hand root position | |
| translation range | |
| translation range |
| Parameter | Setting |
|---|---|
| Simulation timestep | s |
| Gravity | |
| Render size | |
| Controlled DoF | Index, MRS finger group |
| Index joint range | rad |
| MRS joint range | rad |
| Task | Success criterion |
|---|---|
| Reach | |
| Orient | and |
| Shape | and direction agreement as defined below |
| Grasp | and |
| Carry | and |
| Orient2 | and |
| Parameter | Setting |
|---|---|
| Base model | OpenVLA-7B |
| Fine-tuning method | LoRA + action head + proprioceptive projector |
| Optimization steps | 50,000 |
| Effective batch size | 8 |
| Gradient accumulation steps | 1 |
| Optimizer | AdamW |
| Parameter | Setting |
|---|---|
| Base model | OpenVLA-7B |
| Fine-tuning method | LoRA + action head + proprioceptive projector |
| Optimization steps | 50,000 |
| Effective batch size | 8 |
| Gradient accumulation steps | 1 |
| Optimizer | AdamW |
| Setting | Value |
|---|---|
| Model dimension | 256 |
| Latent queries | 38 |
| Encoder latents | 38 ( num_encoder_latents ) |
| Cross-attention | 1 head, head dim 64, RoPE on Q/K/V |
| Self-attention depth | 4 |
| Self-attention heads | 8, head dim 64 (inner dim 512) |
| Parameter | Setting |
|---|---|
| Embedding dimension | 4096 |
| Batch size | 32 |
| Neural pooling | Mean over latent slots |
| Language pooling | Masked mean over valid instruction tokens |
| Embedding normalization | normalization |
| Similarity function | Cosine similarity |
| Segment | Positions | Length | Content and embedding source |
| Text prefix | 1–10 | 10 | <s> In: What action should the robot take to ; standard frozen OpenVLA token embeddings. |
| Neural slots | 11–48 | 38 | One 4096-dimensional embedding per slot, produced from causal spiking activity by the neural bridge. Placeholder lookup embeddings are completely replaced. |
| Text suffix | 49–53 | 5 | ?\n Out:␣ ; standard frozen OpenVLA token embeddings. |
| Policy prompt | 1–53 | Fixed across tasks; no task-language tokens are supplied to the policy. | |
| Action targets | 54–109 | 56 | Eight future actions seven action dimensions, represented using OpenVLA action tokens during training. |
| Stop token | 110 | 1 | End-of-sequence supervision used during action training. |
| Policy | D1 ( ) | D1 (SR) | D2 ( ) | D2 (SR) |
|---|---|---|---|---|
| OpenVLA-7B + task heads | 0.86 0.03 | 0.85 0.06 | 0.39 0.02 | 0.99 0.01 |
| OpenVLA-7B + LoRA + task heads | 0.89 0.02 | 0.97 0.04 | 0.49 0.02 | 1.00 0.00 |