Autonomous robots on one-shot missions run over horizons far longer than the trajectories seen during training, under a fixed onboard compute budget. We present FRANK, a 507K-parameter recurrent architecture that combines tau-gated recurrent modules, content-addressable memory, and a feedforward reflex pathway. We evaluate it against recurrent, state-space, and reduced modular baselines at matched parameter count on four algorithmic sequence tasks, trained at length 5-20 and evaluated out to two million tokens. At 100,000x the maximum training length, 6 of 10 FRANK seeds retain exactly 100.0% accuracy, while none of the 50 baseline configurations does, five architectures at ten seeds each with none left incomplete (Fisher exact, two-sided p = 4.2E-6. Targeted lesions across the four tasks yield four distinct component-reliance profiles, consistent with task-dependent allocation across the recurrent, memory, and reflex pathways. Separately, a FRANK policy trained in simulation drives a physical ground vehicle to commanded waypoints through obstacles without teleoperation.
Figures & tables
Fig. 1: System overview. (a) The deployed vehicle. (b) The FRANK policy runs onboard a Raspberry Pi 5 on a ROSMASTER X3, consuming 13 RPLIDAR A1 rays and the body-frame waypoint offset and emitting wheel velocities at 50 Hz; a watchdog zeroes the command if lidar scans stop.
Fig. 2: The FRANK architecture, after [ 1 ] . All components active every timestep; no routing. The inhibit signal blends the recurrent and reflex output paths.
Fig. 3: (a) Per-seed FRANK accuracy against evaluation scale on the state-transition task; six seeds are exactly flat at 100% out to two million tokens. (b) Seeds reaching perfect accuracy at 105× per architecture; the best baseline seed at that scale reaches 30.3% (GRU).
Fig. 4: Targeted lesion at 10% damage. Four tasks, four distinct reliance profiles from the same architecture, with no routing, no task labels, and no component-specific loss.
Fig. 5: Navigation in simulation. (a) The obstacle field used for training. (b) Commanded waypoint (green) and the executed path. (c) The 13 lidar rays that form the policy’s entire exteroceptive observation.