cs.ROJun 27, 2026

J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor Graphs

Authors: Guanqun CaoLiang Chen

Abstract

Classical simultaneous localization and mapping (SLAM) estimates metric poses and a geometric map but does not provide an action-conditioned predictive state. Action-conditioned world models learn compact latent dynamics but ignore global metric consistency and accumulate drift under open-loop rollout. We introduce J-LAW (Joint Localization and Action-Conditioned World Modeling), a unified factor-graph formulation that connects metric pose variables, predictive latent states, and persistent latent landmarks in this letter.J-LAW represents each image as a compact predictive state and combines it with pose or motion measurements through a separately learned mapping. Its maximum a posteriori (MAP) factor graph enforces consistency between these complementary sources of information over time. Experiments on PushT and WildGS show that J-LAW's factor-graph representation can improve long-horizon latent consistency and recover more reliable predictive states under partial observations, forming a foundation for future integrated localization and planning systems.

Explore similar work

CardsList
  1. When Does LeJEPA Learn a World Model?

    May 25, 2026David Klindt, Yann LeCun, Randall BalestrieroLatent World ModelsWorld Models

  2. D-JEPA: A Decision-Aligned Latent World Model

    Sep 21, 2026Shuaijun Liu, Chengyu Wu, Qifu Wen +5Latent World Models