Papers
arxiv:2609.33804

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

Published on Sep 27
· Submitted by
Louie Hong Yao
on Oct 5
Authors:
,
,
,
,
,
,

Abstract

Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.

Community

Paper submitter

From microscopic molecular motion to macroscopic video scenes, a key challenge in building physical intelligence is understanding spatial structure in dynamic systems and how it evolves over time — in other words, spatiotemporal perception.

Traditional approaches often describe temporal evolution through recurrent structures, differential equations, or specific physical mechanisms. Transformers offer a more flexible, data-driven alternative, but the relationship between space and time is typically left for the model to learn implicitly from data. This naturally raises a question: can we preserve the flexibility of data-driven learning while introducing a more explicit geometric prior for joint spatiotemporal modeling?

To this end, we propose MinkowskiPE. It borrows the algebraic structure of Minkowski geometry and Lorentz transformations, using joint spacetime coordinates to transform the query and key features. Although each token is encoded independently using its own absolute coordinate, the resulting attention score depends only on the relative spacetime displacement between the two tokens. This gives MinkowskiPE spacetime translation invariance while preserving the standard attention computation, making it directly compatible with efficient implementations such as FlashAttention.

Importantly, using Minkowski geometry does not assume that the underlying physical systems obey relativistic dynamics. Instead, it serves as a geometric inductive bias at the representation level. We evaluate the same positional encoding formulation on two very different scales — protein–ligand molecular dynamics and KTH video prediction — and observe consistent performance improvements on both tasks.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33804
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33804 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33804 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33804 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.