UniQuery4R / README.md
KosmoResearch's picture
Update README.md
f1143e4 verified
|
Raw History Blame Contribute Delete
2.8 kB
---
license: apache-2.0
library_name: pytorch
pipeline_tag: image-to-3d
tags:
- 4d-reconstruction
- dynamic-scene
- scene-flow
- point-cloud
- point-tracking
- query-based
---
# UniQuery4R: Unified 4D Scene Reconstruction from a Single Query
[Paper](https://arxiv.org/abs/2608.17283) | [Project Page](https://kosmoresearch.github.io/UniQuery4R/) | [Code](https://github.com/kosmoresearch/UniQuery4R)
Tiancheng Chen<sup>1</sup>, Sheng Tang<sup>1</sup>, Wenhua Jin<sup>1</sup><sup>,</sup><sup>2</sup>, Weiqi Zhang<sup>3</sup>, Juntong Fang<sup>3</sup>, Junsheng Zhou<sup>3</sup>, Zesong Li<sup>1</sup>
<sup>1</sup>Kosmo Research &nbsp;&nbsp; <sup>2</sup>Automotive Engineering Department, Jilin University &nbsp;&nbsp; <sup>3</sup>School of Software, Tsinghua University
**UniQuery4R** is a query-conditioned feed-forward framework for unified 4D scene reconstruction from a single query. A variable-length multi-view clip is jointly encoded **once**; at decoding time, a continuous source-pixel query `q = (u, v, src, tgt)` selects the source and target views via source-to-target cross-attention. Each query jointly predicts target correspondence (`warp2d`), target-time 3D position (`warp3d`), scene flow (`warp3d_delta`, parameterized as direction x magnitude) and source depth, while camera parameters are estimated per view.
## Model files
| File | Description |
| --- | --- |
| `uniquery4r.pth` | UniQuery4R weights (ViT-g encoder, ~1.65B parameters) |
## Usage
Install the code from the [GitHub repository](https://github.com/kosmoresearch/UniQuery4R), then either download this file manually or let `huggingface_hub` fetch it automatically by passing the repo id:
```bash
python scripts/infer_4d.py --checkpoint Kosmo-Research/UniQuery4R --input path/to/images --output outputs/demo
```
```python
from uniquery4r.models.uniquery4r import UniQuery4R
model = UniQuery4R.from_pretrained("Kosmo-Research/UniQuery4R", device="cuda")
import torch
images = ... # [S, 3, H, W] float tensor in [0, 1]
with torch.inference_mode():
predictions = model(images)
```
See the [GitHub README](https://github.com/kosmoresearch/UniQuery4R) for the full CLI, motion-mask, viewer and WorldTrack-evaluation workflows.
## Citation
```bibtex
@article{chen2026uniquery4r,
title={{UniQuery4R}: Unified {4D} Scene Reconstruction from a Single Query},
author={Chen, Tiancheng and Tang, Sheng and Jin, Wenhua and Zhang, Weiqi and Fang, Juntong and Zhou, Junsheng and Li, Zesong},
journal={arXiv preprint arXiv:2608.17283},
year={2026}
}
```
## License
Released under the [Apache License 2.0](LICENSE). Portions of the codebase are adapted from [VGGT](https://github.com/facebookresearch/vggt) and remain under the VGGT License (non-commercial); see [NOTICE](NOTICE) for details.