LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
Abstract
Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.
Community
Whose “left” is it—the camera’s, another person’s, or an imagined observer’s? VLMs often mix these up. We introduce LeRF, which learns to make the requested viewpoint explicit before reasoning.
When needed, the model locates the reference person or object and predicts its front, left, and up directions. A lightweight tool draws these axes on the image, and the model uses the visual cues to answer the question. Supervised fine-tuning teaches frame prediction and selective tool use; reinforcement learning teaches reasoning with these self-generated frames.
No external perception models or explicit 3D reconstruction are needed at inference time. LeRF improves perspective-taking performance across multiple benchmarks, including questions about imagined viewpoints.
We’d love to hear your thoughts, questions, and ideas—especially on how learned reference frames could support broader spatial reasoning and embodied AI!
Get this paper in your agent:
hf papers read 2609.36219 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
SamuelBang/LeRF-4B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper