GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
Abstract
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by 12.7% and 23.1% on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
Community
We introduce GAE — Geometry-Native Autoencoder, which enables video generation directly in a geometry-native latent space. Built from geometry foundation model features, this compact representation serves as a shared space for perception and generation. With camera trajectories as conditioning, the video model generates a unified state that jointly decodes into RGB, depth, and 3D geometry.
Project page: https://jiah-cloud.github.io/GAE.github.io/
GitHub: https://github.com/TencentARC/GAE-GeometricAutoEncoder
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- Beyond Pixels: From Video Priors to 4D Worlds (2026)
- RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation (2026)
- UniSpace: Unified Visual Representation and Scalable Multimodal Modeling (2026)
- GenRec: Knowing Where to Reconstruct and Where to Generate (2026)
- USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes (2026)
- SpatialCrafter: Single Image World Modeling with Generative 3D Proxies (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.24981 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
