Papers
arxiv:2609.36219

LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

Published on Sep 28
· Submitted by
Bang Xiao
on Sep 29
Authors:
,
,
,
,
,
,
,

Abstract

Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.

Community

Paper submitter

Whose “left” is it—the camera’s, another person’s, or an imagined observer’s? VLMs often mix these up. We introduce LeRF, which learns to make the requested viewpoint explicit before reasoning.

When needed, the model locates the reference person or object and predicts its front, left, and up directions. A lightweight tool draws these axes on the image, and the model uses the visual cues to answer the question. Supervised fine-tuning teaches frame prediction and selective tool use; reinforcement learning teaches reasoning with these self-generated frames.

No external perception models or explicit 3D reconstruction are needed at inference time. LeRF improves perspective-taking performance across multiple benchmarks, including questions about imagined viewpoints.

We’d love to hear your thoughts, questions, and ideas—especially on how learned reference frames could support broader spatial reasoning and embodied AI!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36219
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36219 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36219 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.