Compared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data. Neural Robot Dynamics (NeRD) pursues this by keeping collision detection analytical and replacing a simulator's numerical dynamics for the robot with a learned model. Three limitations remain. NeRD offers little speedup over the simulator it learned from, accepts contact only at predefined points, and has no two-way coupling with objects it manipulates. FlashNeRD removes all three with a parallel streaming architecture that makes each prediction faster and more accurate, an encoder that accepts contacts wherever they occur, and two-way coupling with objects simulated by analytical solvers. Experiments across five robots show faster and more accurate dynamics, faster policy learning, and faster inference-time planning. Across three robots, FlashNeRD's dynamics model is up to 55× faster than an optimized NeRD and more accurate over long rollouts. This speedup extends to policy learning, where PPO trains an ANYmal locomotion policy in 34 seconds, 3.8× faster than the analytical simulator and 2.7× faster than an optimized NeRD. With DIAL-MPC, a sampling-based MPC method, the robot climbs all six test platforms where fixed-contact NeRD manages one, at approximately half the analytical simulator's planning time. Cube-reorientation policies trained with FlashNeRD complete within 2% of the simulator-trained policy's target count.
Figures & tables
Fig. 1: FlashNeRD overview (variable-contact case). (a) A shared row encoder and learned pooling compress detected contacts into M fixed-size summary slots plus a count, accepting contacts anywhere on the robot. (b) One of L blocks. Attention reads the current frame and a short cache of past keys and values, while a gated feedforward branch transforms the current frame. Both run in parallel from one packed projection, processing only the newest frame per step. (c) Readout. Temporal features, the current frame, and a direct affine path predict the state increment, retaining current-state information. Green outlines mark learned modules, and green arrows carry contact information.
Fig. 2: Bounded KV reuse. (a) Each block attends over a short span of cached key/value pairs. Because deeper blocks read hidden states that already summarize earlier frames, the receptive field in raw frames (five here, from two blocks of span three) exceeds any single span. Bounding the span keeps streaming inference identical to training. (b) Each step reads the cached pairs and the new pair Pi=(ki,vi) (bold outline) in lag order, then overwrites the oldest slot, so the buffer is never shifted.
Fig. 3: Two-way coupling. FlashNeRD predicts free robot motion. MJWarp supplies analytical object dynamics, solves contacts and joint constraints for robot and objects jointly, and integrates both. Object reactions act on the robot, whose resolved state feeds the next prediction.
Fig. 4: Model-only throughput and prediction accuracy on CartPole, Franka, and G1. Each spoke shows reciprocal rollout NRMSE AH=1/EH at H=1,15,100,200 or throughput QB at batch sizes 16, 16,384, and 512,000. Values are on a logarithmic radial scale normalized by the better model on each spoke, so a larger radius is better. Rim labels give the FlashNeRD-to-Optimized-NeRD ratio, colored by the favored method.
System
Model
E1
E15
E100
E200
CartPole
Opt. NeRD
5.89×10−5
4.78×10−4
8.43×10−3
0.106
FlashNeRD
4.41×10−5
1.73×10−4
3.23×10−3
0.0599
Franka
Opt. NeRD
5.005×10−4
5.63×10−3
0.291
0.747
FlashNeRD
5.008×10−4
5.00×10−3
0.232
0.674
G1
Opt. NeRD
0.0351
0.0847
0.247
0.319
FlashNeRD
0.0231
0.0667
0.230
0.302
TABLE I: Recursive-rollout NRMSE and model-only throughput.
Fig. 5: PPO training reward versus accumulated training time (sum of logged iteration durations) for ANYmal-C velocity tracking with 16,384 environments over 300 iterations, with fixed contact for all learned models. Diamonds mark final policies evaluated in MJWarp with mean actions and the training reward, showing faster training reaches comparable quality.
Fig. 6: Microduck forward rolling with PPO fine-tuning in Newton and FlashNeRD, both evaluated in Newton. Bars show mean wall time for 100 PPO iterations over two seeds.
Fig. 7: DIAL-MPC climbing a 0.181 m platform. Rows are the planning dynamics and columns are matched times. All selected actions are executed in MJWarp within Newton, so only the planner’s model differs. The robot climbs with MJWarp and with variable-contact FlashNeRD, and fails to clear the edge with fixed-contact NeRD.
Rollout dynamics
Climbs
Planning time (ms)
NeRD with fixed contact
1/6
2441
MJWarp
6/6
3528
FlashNeRD with variable contact
6/6
1720
TABLE II: DIAL-MPC on six terrains with execution in Newton (MJWarp). Planning time is the median per control decision.
Fig. 8: Cube reorientation in MJWarp for policies trained with different dynamics (rows). Small cubes show target orientations and counts show completions. The policy trained with one-way coupling never felt the cube resist during training and drops it (arrow), while the two-way FlashNeRD and MJWarp-trained policies complete four targets in the same time.
As robot control shifts toward large-scale reinforcement learning with in-loop dynamics computation, the community's reliance on CPU-bound libraries such as Pinocchio creates a throughput bottleneck in GPU-based training pipelines. We present BARD (Batched Articulated Rigid-body Dynamics), a self-contained PyTorch implementation of Featherstone's rigid-body dynamics algorithms, optimized for batched GPU evaluation and automatic differentiation. Three design choices make this efficient: a tiered lazy-evaluation cache that avoids redundant tree traversals, matmul-free joint transforms via pre-computed Rodrigues constants, and level-parallel propagation that reduces sequential operations to tree-depth batched steps. On five robot models (7-23 DOFs), BARD matches Pinocchio numerically while reaching up to 64x higher throughput for Forward Kinematics and 63x for Jacobians at batch size 4096 on an NVIDIA H200. We validate differentiability through gradient-based system identification on a 7-DOF manipulator, recovering link masses to 1.24% mean error under 5% torque noise, and integrate BARD into an Isaac Lab AMP training pipeline for an 11-DOF spined quadruped with 4096 parallel environments, where it is 8.5x faster than Pinocchio and 2.0x faster than ADAM for in-loop dynamics. BARD is open-sourced at: https://github.com/YueWang996/bard-pytorch-dynamics.
Yue Wang, Yanran Xu, Wenbo Wu +2
University of Southampton, Southampton, SO17 1BJ, UK
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates (≥92% across all tasks), a per-episode inference time of 31.40ms (up to 175× faster than diffusion policies and 18× faster than prior flow matching policies), up to 4× faster training convergence than ACT, and 5× to 7× reduction in controller tracking error compared to discrete-action baselines.
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
MARS Lab, Nanyang Technological University, Singapore
Training Deep Reinforcement Learning (DRL) navigation policies for different robot configurations remains time-consuming. We present FlashNav, a GPU-based framework that trains robot-specific navigation policies within tens of seconds. A unified robot specification configures a lightweight simulator for batched motion updates, range sensing, and footprint collision checking over a shared occupancy map. The framework supports nonconvex footprints, different sensor configurations and drive types. Blockwise ray queries and selective observation recomputation after resets reduce simulation overhead, while GPU-resident replay and overlapping experience collection and learner updates support efficient off-policy training. Experiments covered five robot configurations and three computing platforms. With FastDSAC on a single RTX 5090 GPU, FlashNav can train a deployable navigation policy in under 30 seconds. FlashNav achieved the highest success rate and score in the benchmark comparison. The selected policies were deployed on wheeled, quadrupedal, humanoid, and irregularly shaped robots without additional policy training.
Shanze Wang, Yiwei Qian, Xinming Zhang +8
Eastern Institute of Technology, Ningbo · The Hong Kong Polytechnic University · National University of Singapore +2