Compared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data. Neural Robot Dynamics (NeRD) pursues this by keeping collision detection analytical and replacing a simulator's numerical dynamics for the robot with a learned model. Three limitations remain. NeRD offers little speedup over the simulator it learned from, accepts contact only at predefined points, and has no two-way coupling with objects it manipulates. FlashNeRD removes all three with a parallel streaming architecture that makes each prediction faster and more accurate, an encoder that accepts contacts wherever they occur, and two-way coupling with objects simulated by analytical solvers. Experiments across five robots show faster and more accurate dynamics, faster policy learning, and faster inference-time planning. Across three robots, FlashNeRD's dynamics model is up to 55× faster than an optimized NeRD and more accurate over long rollouts. This speedup extends to policy learning, where PPO trains an ANYmal locomotion policy in 34 seconds, 3.8× faster than the analytical simulator and 2.7× faster than an optimized NeRD. With DIAL-MPC, a sampling-based MPC method, the robot climbs all six test platforms where fixed-contact NeRD manages one, at approximately half the analytical simulator's planning time. Cube-reorientation policies trained with FlashNeRD complete within 2% of the simulator-trained policy's target count.
Figures & tables
Fig. 1: FlashNeRD overview (variable-contact case). (a) A shared row encoder and learned pooling compress detected contacts into M fixed-size summary slots plus a count, accepting contacts anywhere on the robot. (b) One of L blocks. Attention reads the current frame and a short cache of past keys and values, while a gated feedforward branch transforms the current frame. Both run in parallel from one packed projection, processing only the newest frame per step. (c) Readout. Temporal features, the current frame, and a direct affine path predict the state increment, retaining current-state information. Green outlines mark learned modules, and green arrows carry contact information.
Fig. 2: Bounded KV reuse. (a) Each block attends over a short span of cached key/value pairs. Because deeper blocks read hidden states that already summarize earlier frames, the receptive field in raw frames (five here, from two blocks of span three) exceeds any single span. Bounding the span keeps streaming inference identical to training. (b) Each step reads the cached pairs and the new pair Pi=(ki,vi) (bold outline) in lag order, then overwrites the oldest slot, so the buffer is never shifted.
Fig. 3: Two-way coupling. FlashNeRD predicts free robot motion. MJWarp supplies analytical object dynamics, solves contacts and joint constraints for robot and objects jointly, and integrates both. Object reactions act on the robot, whose resolved state feeds the next prediction.
Fig. 4: Model-only throughput and prediction accuracy on CartPole, Franka, and G1. Each spoke shows reciprocal rollout NRMSE AH=1/EH at H=1,15,100,200 or throughput QB at batch sizes 16, 16,384, and 512,000. Values are on a logarithmic radial scale normalized by the better model on each spoke, so a larger radius is better. Rim labels give the FlashNeRD-to-Optimized-NeRD ratio, colored by the favored method.
System
Model
E1
E15
E100
E200
CartPole
Opt. NeRD
5.89×10−5
4.78×10−4
8.43×10−3
0.106
FlashNeRD
4.41×10−5
1.73×10−4
3.23×10−3
0.0599
Franka
Opt. NeRD
5.005×10−4
5.63×10−3
0.291
0.747
FlashNeRD
5.008×10−4
5.00×10−3
0.232
0.674
G1
Opt. NeRD
0.0351
0.0847
0.247
0.319
FlashNeRD
0.0231
0.0667
0.230
0.302
TABLE I: Recursive-rollout NRMSE and model-only throughput.
Fig. 5: PPO training reward versus accumulated training time (sum of logged iteration durations) for ANYmal-C velocity tracking with 16,384 environments over 300 iterations, with fixed contact for all learned models. Diamonds mark final policies evaluated in MJWarp with mean actions and the training reward, showing faster training reaches comparable quality.
Fig. 6: Microduck forward rolling with PPO fine-tuning in Newton and FlashNeRD, both evaluated in Newton. Bars show mean wall time for 100 PPO iterations over two seeds.
Fig. 7: DIAL-MPC climbing a 0.181 m platform. Rows are the planning dynamics and columns are matched times. All selected actions are executed in MJWarp within Newton, so only the planner’s model differs. The robot climbs with MJWarp and with variable-contact FlashNeRD, and fails to clear the edge with fixed-contact NeRD.
Rollout dynamics
Climbs
Planning time (ms)
NeRD with fixed contact
1/6
2441
MJWarp
6/6
3528
FlashNeRD with variable contact
6/6
1720
TABLE II: DIAL-MPC on six terrains with execution in Newton (MJWarp). Planning time is the median per control decision.
Fig. 8: Cube reorientation in MJWarp for policies trained with different dynamics (rows). Small cubes show target orientations and counts show completions. The policy trained with one-way coupling never felt the cube resist during training and drops it (arrow), while the two-way FlashNeRD and MJWarp-trained policies complete four targets in the same time.