LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
Authors: Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg
Organizations: University of Illinois Urbana-Champaign · Zhiyuan College, Shanghai Jiao Tong University · University of California, Merced · Shanghai Jiao Tong University · Amazon AGI
Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.
Figures & tables
Figure 1 : Motivation of LeRF. Perspective taking requires reasoning with respect to the viewpoint specified by a spatial query. An explicit entity-centered reference frame makes this viewpoint visually accessible, helping VLMs resolve viewpoint-dependent spatial relations and motivating LeRF to learn and reason with self-grounded reference frames.
Figure 2 : Overview of LeRF. The VLM selectively constructs and renders a reference frame for perspective-taking reasoning, or directly reasons from the original image when no frame is needed. LeRF is trained with SFT for frame prediction and RL for reasoning with self-predicted frames.
Method
OmniSpatial-PT
3DSRBench
ViewSpatial-Bench
Ego
Allo
Hypo
Ori
M-Obj
P-Obj
P-Rel
Proprietary Model
GPT-5.6-Luna (medium)
83.33
49.73
45.78
60.04
55.34
46.99
70.07
GPT-5.6-Terra (medium)
81.37
55.85
53.01
63.32
56.39
45.08
77.20
Claude Sonnet 5 (medium)
80.39
42.55
49.40
34.94
44.77
51.31
51.43
Claude Sonnet 5 (high)
84.31
48.14
45.78
43.15
46.81
51.51
60.10
Table 1 : Quantitative results on OmniSpatial-Perspective Taking [ 9 ] (OmniSpatial-PT), 3DSRBench [ 16 ] , and ViewSpatial-Bench (V-Spatial-Bench) [ 14 ] . Ori / M-Obj denote Orientation and Multi-Object, and P-Obj / P-Rel denote Person Perspective–Object View Orientation and Person Perspective–Relative Direction, respectively. Except for SpatialReasoner, which does not support thinking, all models enable thinking during evaluation. The best and second-best results among open-source methods are shown in bold and underlined .
Model
EMDB
OmniNOCS-Objectron
All@15 ∘
All@30 ∘
Ori@0.1Diag
All@15 ∘
All@30 ∘
Ori@0.1Diag
Baseline VLMs
Qwen3.5-4B
6.87
17.89
20.45
2.73
5.66
53.04
Qwen3.5-9B
1.11
4.43
91.63
0.52
1.68
70.34
Expert Model
OriAnything-V2
19.35
44.84
–
50.42
83.23
–
Table 2 : Projected reference-frame accuracy (%) on sampled EMDB and OmniNOCS-Objectron subsets. All@ θ requires all three 2D axis-direction errors to be within θ . Origin@0.1Diag denotes the percentage of views whose predicted origin is within 10% of the ground-truth bounding-box diagonal from the ground-truth origin. Best and second-best results are bold and underlined .
Figure 5
RL Initialization
RL Render
Eval Render
Ego
Allo
Hypo
Overall
Base Model
Qwen3.5-9B (Base)
–
–
80.20
47.13
44.58
52.76
SFT Baseline
Qwen3.5-9B + SFT
–
✓
69.80
34.47
34.22
40.86
Direct RL Fine-tuning (w/o SFT)
Qwen3.5-9B + RL
✗
✗
80.34
49.40
45.18
54.34
Table 3 : Effect of SFT, RL, and reference-frame rendering on OmniSpatial-PT. The base model and SFT-only checkpoint are included as baselines. All RL variants are trained for 50 steps. When rendering is disabled, the tool returns the original image. Accuracy is reported in percentages.
Figure 5 : Qualitative examples of LeRF-9B on egocentric, allocentric, and hypothetical perspective-taking queries.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Effect of coordinate-axis length on accuracy.
Table 5 : Inference efficiency comparison. All models are measured under single-request inference on an RTX PRO 6000 GPU.
Figure 7 : Additional qualitative comparison with SpatialReasoner, APC+Qwen3.5-9B, and Qwen3.5-9B. Baseline models frequently answer according to camera-centric spatial relations or fail to resolve the requested viewpoint, whereas LeRF explicitly constructs an entity-centered reference frame and correctly reasons from the specified perspective.
Figure 8 : Representative failure cases of LeRF. Errors can arise from inaccurate reference-frame prediction, incorrect grounding of the reference entity, or downstream reasoning failures even when the reference frame is reasonably constructed.
Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understanding, such as viewpoint reasoning, directional comparison, and distance estimation. In multi-view images and monocular videos, relevant spatial cues are often sparse and distributed across redundant observations, making them difficult to organize and exploit. Reconstruction-based Vision Foundation Models (VFMs) offer a natural way to aggregate such observations into explicit spatial memory, such as point clouds. However, simply exposing reconstruction models as free-form tools is brittle, VLMs may invoke tools incorrectly, skip required spatial transformations, or misuse intermediate results. We propose \textbf{Reasmory}, a framework that formulates spatial reasoning as structured program execution over reconstructed spatial memory. Reasmory constructs explicit 3D memory, augments it with semantically grounded 3D object instances, and introduces a lightweight Domain-Specific Language (DSL) that constrains how VLMs query objects and cameras, transform viewpoints, and render observations during reasoning. Generated programs are parsed and validated before execution, enabling more reliable interaction with spatial memory than unconstrained tool use. Experiments on multi-view image and video spatial reasoning benchmarks show consistent gains of 6--18% over strong baselines, including GPT-5-mini and Gemini-3-flash, indicating that explicit 3D memory is most useful when accessed through constrained, validated operations rather than free-form tool calls.
Jixuan He, Xueting Li, Chieh Hubert Lin +1
Cornell Tech, Cornell University · NVIDIA · illoca AI +1
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path Spatial Reasoning via Reinforcement Learning for Spatial VLMs (SR-REAL), a unified framework that equips a spatial VLM with two complementary reasoning paths: Language-Only Reasoning (LOR), which performs step-by-step linguistic deduction, and Detect-Then-Reason (DTR), which detects 3D geometric cues (e.g., centers or bounding boxes) via region tokens before explicit geometric inference. SR-REAL begins with a cold-start supervised fine-tuning stage that constructs LOR and DTR chain-of-thought supervision and exposes a region-to-3D interface, followed by RL that optimizes the policy model with accuracy and format rewards; for DTR, a discrete center-based detection reward further refines geometric alignment. Across diverse spatial benchmarks, SR-REAL significantly outperforms spatial VLM baselines: (i) a single RL-trained model supports both reasoning paths, with DTR excelling in region-aware tasks through precise 3D localization and LOR enhancing general spatial reasoning; (ii) jointly training both paths fosters mutual reinforcement; (iii) high-quality, blended cold-start data is crucial for stable RL optimization; and (iv) the model generalizes across datasets and domains without per-task tuning, demonstrating positive transfer between LOR and DTR.
Yatai Ji, An-Chieh Cheng, Yang Fu +13
The University of Hong Kong · NVIDIA · University of California, San Diego
Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognition. Existing methods typically build such representations through textual reasoning or 3D reconstruction. However, they often falter during multi-step reasoning, particularly when required to dynamically re-anchor evidence to the specific camera-, object-, or direction-centric reference frames demanded by complex queries. To address this, we propose OmniView-Space, a framework designed to maintain spatial consistency through multimodal egocentric evidence. Our approach consists of three core components: (1) Multi-Perspective Spatial Mapping (MPSM), which re-anchors reconstructed geometry into a query-aligned visual cognitive map and a textual spatial graph; (2) Tool-Guided Egocentric Reasoning, an interleaved policy trained to actively select the ego anchor required by the query and request the corresponding MPSM evidence; and (3) Cognitive-Map Distillation, which uses MPSM-generated trajectories and ego-frame rewards to train the model to reason with self-generated cognitive maps. Experiments on single- and multi-image spatial reasoning benchmarks show that OmniView-Space achieves state-of-the-art performance. Furthermore, the distilled model maintains this performance while reducing reliance on external geometry pipelines.
Xudong Li, Mengdan Zhang, Peixian Chen +7
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China · Tencent Youtu Lab · Beijing Institute of Technology