A Biophysically Detailed C. elegans Circuit as a Task-Agnostic Dynamical Core for Visually Robust Robot Manipulation
Authors: Linrui Qian, Jiajia Zhang, Gan He, Bohan Sun, Zhiwei Lin, Qianhao Wang, Zewu Cai, Nianyu Yi, +2 more
Organizations: CogLeap.AI Space Intelligence (Wuxi) Technology Co., Ltd., Beijing 100080, China · School of Mathematics and Computational Science, Xiangtan University, Xiangtan 411105, China · Institute for Brain and Intelligence, Fudan University, Shanghai 200433, China. · Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing 100084, China
Robot policies are usually trained for one task, one body and one visual environment, and generalize poorly beyond these conditions. Whether a nervous system can instead supply the sensorimotor computation through its evolved wiring and biophysics remains unresolved. Here we embed a biophysically detailed Caenorhabditis elegans sensorimotor circuit - 136 multicompartment neurons with realistic morphologies and electrophysiological characteristics - as the dynamical core of a visuomotor policy. Only thin task-specific adapters are trained; the core's synaptic weights stay fixed while its membrane voltages evolve freely. Across different MetaWorld tasks the core matches or exceeds diffusion-policy, action-chunking-transformer and neural-circuit-policy baselines, and degrades less under visual perturbations. Replacing the core with generic network models such as MLP, LSTM, transformer or reservoir networks removes the advantage. Furthermore, on a real robotic arm the core withstands diverse visual perturbations that collapse the baselines. Our results suggest that visual robustness can be inherited from biophysically detailed circuit dynamics rather than learned by task-specific controllers.
Figures & tables
Figure 1: A biophysically detailed C. elegans sensorimotor circuit as the task-agnostic dynamical core of a visuomotor policy. Continuous RGB frames are compressed into a low-dimensional sensory drive that is injected into the connectome-constrained circuit as clamp currents. The circuit’s motor-neuron state is conditioned with visual features by FiLM and decoded into an action chunk that is executed in a receding-horizon loop (predict eight steps, execute four, re-observe and replan). The circuit contains 136 multicompartment neurons and 1,901 interneuron connections; its parameters are fixed after training and its membrane voltages evolve freely, and only the encoder, FiLM parameters and action decoder are trained. The circuit is the complete sensorimotor model of BAAIWorm ( Zhao et al., 2024 ) , reused here without modification.
Figure 2: A single core across four simulated manipulation tasks and three perturbations. a , Simulator frames of the four tasks (coffee-push, hammer, door-open, assembly). b , Clean success rate of the core, NCP, ACT and DP; the core matches or exceeds every baseline. c , The perturbations, shown on simulator frames: additive image noise ( σ=0.10 ), lighting change, background change and object displacement. d , Success rate under each perturbation, one panel per task; within each tick the bars are Ours, NCP, ACT and DP.
Figure 3: Robustness is carried by the circuit, not by the interface. a , Motor-state trajectories of the full circuit and five replacement cores at three noise levels, projected onto a manifold fitted to the clean trajectories of each model. Clean trajectories are black; colored lines are individual noisy trials. b , Speed profiles on the manifold for moderate noise ( σ=0.15 ), with the band width between the noisy and clean profiles reported above each panel. c , Five dynamical metrics (temporal jitter, normalized jerk, path curvature, velocity dissimilarity and tangling) as a function of the noise level, normalized so that 1 marks the clean reference. All variants share the same encoder, FiLM mechanism, decoder and action representation; only the core differs.
Figure 4: Robustness, deviation dynamics and ablation on a real robot. a , Experimental setup and the four perturbations (sensor noise, lighting change, background change and object change). b , Success rate of the core, ACT and DP under the nominal condition and under each perturbation (Ours red, ACT grey, DP light grey). c , Manifold of each policy’s own representation (Ours motor neural state, ACT action chunk, DP action prediction) projected on the first two principal components fitted to that model’s clean repeats; annotations give the ratio of noisy to clean trial-to-trial spread and the center shift. d , Ablation analysis by core replacement: success rate of the MLP, LSTM, transformer and reservoir replacements together with the full core (Ours red) under the nominal condition and each perturbation.
Visuomotor policies observe both the task scene and the acting embodiment, allowing embodiment-specific visual cues to influence action prediction. We study this phenomenon as visual embodiment dependence (VED) and show, through cue-conflict interventions across representative policies, that visible robot configuration can become a shortcut to task progress. Rather than eliminating VED, we argue that it should be structured around embodiment information that supports control and generalization. We realize this through embodiment canonicalization in 3D point clouds, replacing the original embodiment with a canonical end-effector representation (CER) that preserves control-relevant geometry while abstracting embodiment-specific morphology. Its editable form further enables configuration-decorrelation augmentation for unfamiliar robot configurations. Experiments show that embodiment canonicalization substantially improves human-to-robot policy transfer without robot demonstrations, while simply removing the embodiment is insufficient without preserving control-relevant geometry. We further find that CER itself can become a configuration shortcut when robot configuration becomes decoupled from task progress; configuration-decorrelation augmentation mitigates this failure mode and restores robust recovery without sacrificing performance on seen configurations. Together, these results show that robust visuomotor learning benefits from structuring, rather than removing, visual embodiment information. Project website: https://tonyfang.net/ved
Hongjie Fang, Yuxuan Lu, Chenxi Wang +8
1Shanghai Jiao Tong University · 2FORTE Lab. · 3Noematrix. +2
VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/
Zhenyang Feng, Unnat Jain
Department of Computer Science, University of California, Irvine
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry. Tactile sensing provides these complementary signals, yet tactile data remain costly to collect and hard to generalize across sensors, robots, and tasks. We introduce OmniTacTune, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction. OmniTacTune uses a two-stage design: it first bootstraps tactile-aware learning from autonomous base-policy rollouts, then learns a lightweight tactile residual policy through online interaction. Extensive experiments show that OmniTacTune generalizes across diverse contact-rich tasks, visual base policies, and tactile representations. Across four real-world contact-rich tasks, it improves visual base policies from 5-40% success to 85-100% within 40-80 minutes, demonstrating an efficient path for adapting tactile feedback to scalable visual robot policies. Project page: https://colinyu1.github.io/omnitactune-site/
Kelin Yu, Haode Zhang, Harish Ravichandar +2
University of Maryland, College Park · Georgia Institute of Technology