File size: 2,801 Bytes
77ff5fd
 
9efecd6
 
 
 
 
 
 
 
 
77ff5fd
9efecd6
 
 
 
 
f1143e4
9efecd6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
license: apache-2.0
library_name: pytorch
pipeline_tag: image-to-3d
tags:
  - 4d-reconstruction
  - dynamic-scene
  - scene-flow
  - point-cloud
  - point-tracking
  - query-based
---

# UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

[Paper](https://arxiv.org/abs/2608.17283) | [Project Page](https://kosmoresearch.github.io/UniQuery4R/) | [Code](https://github.com/kosmoresearch/UniQuery4R)

Tiancheng Chen<sup>1</sup>, Sheng Tang<sup>1</sup>, Wenhua Jin<sup>1</sup><sup>,</sup><sup>2</sup>, Weiqi Zhang<sup>3</sup>, Juntong Fang<sup>3</sup>, Junsheng Zhou<sup>3</sup>, Zesong Li<sup>1</sup>

<sup>1</sup>Kosmo Research &nbsp;&nbsp; <sup>2</sup>Automotive Engineering Department, Jilin University &nbsp;&nbsp; <sup>3</sup>School of Software, Tsinghua University

**UniQuery4R** is a query-conditioned feed-forward framework for unified 4D scene reconstruction from a single query. A variable-length multi-view clip is jointly encoded **once**; at decoding time, a continuous source-pixel query `q = (u, v, src, tgt)` selects the source and target views via source-to-target cross-attention. Each query jointly predicts target correspondence (`warp2d`), target-time 3D position (`warp3d`), scene flow (`warp3d_delta`, parameterized as direction x magnitude) and source depth, while camera parameters are estimated per view.

## Model files

| File | Description |
| --- | --- |
| `uniquery4r.pth` | UniQuery4R weights (ViT-g encoder, ~1.65B parameters) |

## Usage

Install the code from the [GitHub repository](https://github.com/kosmoresearch/UniQuery4R), then either download this file manually or let `huggingface_hub` fetch it automatically by passing the repo id:

```bash
python scripts/infer_4d.py --checkpoint Kosmo-Research/UniQuery4R --input path/to/images --output outputs/demo
```

```python
from uniquery4r.models.uniquery4r import UniQuery4R

model = UniQuery4R.from_pretrained("Kosmo-Research/UniQuery4R", device="cuda")

import torch
images = ...  # [S, 3, H, W] float tensor in [0, 1]
with torch.inference_mode():
    predictions = model(images)
```

See the [GitHub README](https://github.com/kosmoresearch/UniQuery4R) for the full CLI, motion-mask, viewer and WorldTrack-evaluation workflows.

## Citation

```bibtex
@article{chen2026uniquery4r,
  title={{UniQuery4R}: Unified {4D} Scene Reconstruction from a Single Query},
  author={Chen, Tiancheng and Tang, Sheng and Jin, Wenhua and Zhang, Weiqi and Fang, Juntong and Zhou, Junsheng and Li, Zesong},
  journal={arXiv preprint arXiv:2608.17283},
  year={2026}
}
```

## License

Released under the [Apache License 2.0](LICENSE). Portions of the codebase are adapted from [VGGT](https://github.com/facebookresearch/vggt) and remain under the VGGT License (non-commercial); see [NOTICE](NOTICE) for details.