cs.CVSep 28, 2026

LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

Authors: Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg

Organizations: University of Illinois Urbana-Champaign · Zhiyuan College, Shanghai Jiao Tong University · University of California, Merced · Shanghai Jiao Tong University · Amazon AGI

Abstract

Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

    May 31, 2026Jixuan He, Xueting Li, Chieh Hubert Lin +1Spatial ReasoningSpatial Memory

  2. Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

    Jun 16, 2026Yatai Ji, An-Chieh Cheng, Yang Fu +13Spatial ReasoningIndoor Localization

  3. OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

    Jul 1, 2026Xudong Li, Mengdan Zhang, Peixian Chen +7Spatial ReasoningMultimodal Reasoning