A Biophysically Detailed C. elegans Circuit as a Task-Agnostic Dynamical Core for Visually Robust Robot Manipulation
Authors: Linrui Qian, Jiajia Zhang, Gan He, Bohan Sun, Zhiwei Lin, Qianhao Wang, Zewu Cai, Nianyu Yi, +2 more
Organizations: CogLeap.AI Space Intelligence (Wuxi) Technology Co., Ltd., Beijing 100080, China · School of Mathematics and Computational Science, Xiangtan University, Xiangtan 411105, China · Institute for Brain and Intelligence, Fudan University, Shanghai 200433, China. · Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing 100084, China
Robot policies are usually trained for one task, one body and one visual environment, and generalize poorly beyond these conditions. Whether a nervous system can instead supply the sensorimotor computation through its evolved wiring and biophysics remains unresolved. Here we embed a biophysically detailed Caenorhabditis elegans sensorimotor circuit - 136 multicompartment neurons with realistic morphologies and electrophysiological characteristics - as the dynamical core of a visuomotor policy. Only thin task-specific adapters are trained; the core's synaptic weights stay fixed while its membrane voltages evolve freely. Across different MetaWorld tasks the core matches or exceeds diffusion-policy, action-chunking-transformer and neural-circuit-policy baselines, and degrades less under visual perturbations. Replacing the core with generic network models such as MLP, LSTM, transformer or reservoir networks removes the advantage. Furthermore, on a real robotic arm the core withstands diverse visual perturbations that collapse the baselines. Our results suggest that visual robustness can be inherited from biophysically detailed circuit dynamics rather than learned by task-specific controllers.
Figures & tables
Figure 1: A biophysically detailed C. elegans sensorimotor circuit as the task-agnostic dynamical core of a visuomotor policy. Continuous RGB frames are compressed into a low-dimensional sensory drive that is injected into the connectome-constrained circuit as clamp currents. The circuit’s motor-neuron state is conditioned with visual features by FiLM and decoded into an action chunk that is executed in a receding-horizon loop (predict eight steps, execute four, re-observe and replan). The circuit contains 136 multicompartment neurons and 1,901 interneuron connections; its parameters are fixed after training and its membrane voltages evolve freely, and only the encoder, FiLM parameters and action decoder are trained. The circuit is the complete sensorimotor model of BAAIWorm ( Zhao et al., 2024 ) , reused here without modification.
Figure 2: A single core across four simulated manipulation tasks and three perturbations. a , Simulator frames of the four tasks (coffee-push, hammer, door-open, assembly). b , Clean success rate of the core, NCP, ACT and DP; the core matches or exceeds every baseline. c , The perturbations, shown on simulator frames: additive image noise ( σ=0.10 ), lighting change, background change and object displacement. d , Success rate under each perturbation, one panel per task; within each tick the bars are Ours, NCP, ACT and DP.
Figure 3: Robustness is carried by the circuit, not by the interface. a , Motor-state trajectories of the full circuit and five replacement cores at three noise levels, projected onto a manifold fitted to the clean trajectories of each model. Clean trajectories are black; colored lines are individual noisy trials. b , Speed profiles on the manifold for moderate noise ( σ=0.15 ), with the band width between the noisy and clean profiles reported above each panel. c , Five dynamical metrics (temporal jitter, normalized jerk, path curvature, velocity dissimilarity and tangling) as a function of the noise level, normalized so that 1 marks the clean reference. All variants share the same encoder, FiLM mechanism, decoder and action representation; only the core differs.
Figure 4: Robustness, deviation dynamics and ablation on a real robot. a , Experimental setup and the four perturbations (sensor noise, lighting change, background change and object change). b , Success rate of the core, ACT and DP under the nominal condition and under each perturbation (Ours red, ACT grey, DP light grey). c , Manifold of each policy’s own representation (Ours motor neural state, ACT action chunk, DP action prediction) projected on the first two principal components fitted to that model’s clean repeats; annotations give the ratio of noisy to clean trial-to-trial spread and the center shift. d , Ablation analysis by core replacement: success rate of the MLP, LSTM, transformer and reservoir replacements together with the full core (Ours red) under the nominal condition and each perturbation.