UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking
Authors: Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, +2 more
Organizations: College of Integrated Circuits, Hunan University, Changsha, China · School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China · School of Advanced Technology, Xi’an Jiaotong-Liverpool University, Suzhou, China · Suzhou VSDeep Intelligent Technology Co., Ltd., Suzhou, China
General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at https://tw5775.github.io/UniTrackPLA.
Figures & tables
Fig. 1: Overview of the proposed UniTrackPLA pipeline. The upper policy branch encodes four perspective views with PAE and fuses them with language instructions to predict waypoint chunks. The lower WAC branch predicts action-conditioned future features and evaluates future-state consistency to dynamically continue execution or trigger replanning. Snowflakes denote frozen modules.
Method
STF
DRF
Overall
SR ↑
FR ↑
LR ↓
CR ↓
SR ↑
FR ↑
LR ↓
CR ↓
SR ↑
FR ↑
LR ↓
CR ↓
TrackVLA †
35.00
65.35
27.00
20.00
12.00
56.27
48.00
20.00
23.50
60.81
37.50
20.00
Ours
51.00
75.65
14.00
15.00
19.00
58.83
46.00
16.00
35.00
67.24
30.00
15.50
TABLE I: Quantitative comparison between TrackVLA and UniTrackPLA on the Omni-Tracking evaluation split.
Method
R2R-CE Seen
R2R-CE Unseen
RxR-CE Seen
RxR-CE Unseen
Overall
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
TrackVLA †
15.00
15.00
14.00
13.80
11.00
10.62
12.00
11.66
13.00
12.77
Ours
19.00
18.40
20.00
20.00
19.00
17.93
21.00
20.88
19.75
19.30
TABLE II: Quantitative comparison between TrackVLA and Ours on the Omni-VLN evaluation split.
Fig. 2: Overview of the Omni-Tracking dataset. It contains single-target following and distractor-rich following under arbitrary 360∘ target initialization. Blue and gray boxes denote the referred target and distractor persons, respectively.
Fig. 3: Qualitative comparison between TrackVLA and Ours for closed-loop panoramic person following. The top and bottom examples correspond to STF and DRF, respectively.
Training Data
ADE ↓
FDE ↓
Yaw MAE ↓
EP@0.2m ↑
Sim-only
0.172
0.288
12.78
42.92
Sim + 10 Real Routes
0.046
0.085
6.05
90.00
Sim + 50 Real Routes
0.042
0.078
5.19
90.83
Sim + 76 Real Routes
0.038
0.069
5.11
92.08
TABLE III: Effect of real-world navigation data on held-out route performance.
Configuration
SR ↑
FR ↑
LR ↓
CR ↓
N=1,FoV=360∘
13.33
59.22
31.67
48.33
N=2,FoV=180∘
5.00
56.81
36.67
53.33
N=4,FoV=90∘
13.33
55.82
33.33
41.67
N=4,FoV=180∘
21.67
61.90
21.67
45.00
N=6,FoV=60∘
2.60
48.96
58.33
28.13
N=6,FoV=120∘
3.33
15.15
41.67
46.67
TABLE IV: Comparison of panoramic view configurations on the direction-balanced configuration split.
Components
Task SR ↑
Overall
PAE
WAC
STF
DRF
SR ↑
FR ↑
LR ↓
CR ↓
✗
✗
35.00
12.00
23.50
60.81
37.50
20.00
✓
✗
47.00
18.00
32.50
67.23
32.00
16.00
✗
✓
47.00
10.00
28.50
62.23
32.00
19.00
✓
✓
51.00
19.00
35.00
67.24
30.00
15.50
TABLE V: Ablation study of PAE and WAC.
Fig. 4: Real-world system architecture. An Insta360 X4 captures panoramic observations on the Go2-W robot, while a remote RTX 4090 server performs inference and returns waypoint commands.
Fig. 7: Effect of real-world training data on outdoor navigation. (a) Navigation using Sim-only training and (b) navigation using Sim + 76 Real routes.
Fig. 8: Real-world experiments in diverse indoor environments. (a) Panoramic navigation in an office and (b) navigation and person tracking in confined domestic spaces.