cs.CVSep 29, 2026

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

Authors: Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Sihan Wang

Organizations: Huazhong University of Science and Technology · Jiangxi University of Water Resources and Electric Power · Zhongnan University of Economics and Law

Abstract

Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects

    Jul 19, 2026Yanrong Qin, Xiaoyan Cao, Yao YaoMulti-Object TrackingRe-Identification

  2. Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

    Sep 1, 2026Orcun Cetintas, Guillem Brasó, Tim Meinhardt +1Multi-Object TrackingMonocular Video

  3. CST-WM: A Causally Structured World Model for Embodied Visual Tracking

    Sep 5, 2026Junyi Hu, Shuaihang Yuan, Jiazhao Liang +1Efficient World-Action ModelWorld Models