File size: 5,047 Bytes
e13fbb6
adaa8dc
 
 
 
 
 
 
 
 
 
 
e13fbb6
9da088f
e13fbb6
45e4ccf
a7281b8
efe0e75
 
858d08d
a7281b8
817a29e
efe0e75
e13fbb6
 
504b215
116d901
e13fbb6
 
 
 
 
116d901
 
45e4ccf
 
 
 
 
1b8024f
 
 
e13fbb6
 
45e4ccf
 
e13fbb6
 
 
 
 
 
 
45e4ccf
 
adaa8dc
20a7db9
adaa8dc
 
e13fbb6
adaa8dc
 
e13fbb6
adaa8dc
e13fbb6
adaa8dc
e13fbb6
adaa8dc
e13fbb6
adaa8dc
 
 
 
 
 
e13fbb6
adaa8dc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e13fbb6
adaa8dc
 
 
 
 
e13fbb6
 
adaa8dc
e13fbb6
adaa8dc
 
 
 
 
 
 
e13fbb6
adaa8dc
e13fbb6
 
adaa8dc
 
 
 
 
 
 
e13fbb6
adaa8dc
 
 
 
 
e13fbb6
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
```yaml
license: cc-by-nc-sa-4.0
language:
- en
tags:
- relighting
- multi-view
- generative-transformer
- 3d-vision
- siggraph-asia-2026
- wan2.1
pipeline_tag: image-to-image
```

<div align="center">

  <span style="display: block; font-size: 60px; font-family: 'Times', monospace; font-weight: bold; color: #555; margin-bottom: 10px;">
    πŸ”† RelightFormer
  </span>

  <span style="display: block; font-size: 20px; font-family: 'Times', monospace; font-weight: bold; color: #555;">
    Feed-Forward Generative Transformer for Multiview Object Relighting
  </span>

</div>

<p align="center">
  <img
    src="https://raw.githubusercontent.com/vLAR-group/RelightFormer/main/demo/teaser.jpg"
    alt="RelightFormer teaser showing photorealistic multi-view relighting results"
    width="80%"
  >
</p>

<p align="center">
  <strong>vLAR Group</strong> | <em>SIGGRAPH Asia 2026</em>
</p>

<p align="center">
  <a href="https://github.com/vLAR-group/RelightFormer">
    <img src="https://img.shields.io/badge/GitHub-Repository-181717?logo=github" alt="GitHub">
  </a>
  <a href="https://arxiv.org/abs/2609.07414">
    <img src="https://img.shields.io/badge/arXiv-2609.07414-b31b1b.svg" alt="arXiv">
  </a>
  <a href="https://huggingface.co/datasets/vLAR/LavalObjaverseDataset">
    <img src="https://img.shields.io/badge/πŸ€—-Dataset-yellow" alt="Dataset">
  </a>
  <a href="https://huggingface.co/vLAR/RelightFormer">
    <img src="https://img.shields.io/badge/πŸ€—-Model-yellow" alt="Model">
  </a>
  <a href="#license">
    <img src="https://img.shields.io/badge/License-CC%20BY--NC--SA%204.0-lightgrey.svg" alt="License">
  </a>
</p>


## 🌟 Overview

**RelightFormer** revolutionizes image relighting by replacing traditional, computationally expensive inverse rendering with a feed-forward generative Transformer. By seamlessly injecting target lighting into spatial features and processing multiple views symmetrically, it delivers highly photorealistic results. Trained on the newly introduced, large-scale open-source **Laval-Objaverse Dataset (LOD )**, RelightFormer achieves state-of-the-art quality and remarkable generalization across diverse scenes.

### ✨ Key Features

- 🏹 **Feed-Forward Architecture**: No iterative optimization required, enabling rapid generation.

- 🌟 **Multi-View Consistency**: Coherent and physically plausible relighting across all viewpoints.

- ⚑ **Performant Inference**: Highly optimized and expeditious execution on modern GPUs.

- 🎨 **Competitive Quality**: State-of-the-art, photorealistic relighting results.

---

## πŸš€ Quick Start

You can easily load and run the model using the `diffsynth` library in our [GitHub Repository](https://github.com/vLAR-group/RelightFormer).
We provide two revisions: `main` (RelightFormer) and `post` (RelightFormer-Post, fine-tuned for enhanced quality).

```python
from diffsynth import RelightFormerPipeline

# Use revision='main' for RelightFormer, or 'post' for RelightFormer-Post
pipe = RelightFormerPipeline.from_pretrained(
    "vLAR/RelightFormer", 
    revision="main"
)

# Example inference (adjust inputs according to your specific pipeline API)
# output = pipe(image=..., lighting=..., ...)
```

> πŸ’‘ **For full inference scripts, multi-GPU evaluation, and training code, please visit the **[**Official GitHub Repository**](https://github.com/vLAR-group/RelightFormer)**.**

---

## πŸ“¦ Dataset

This model is trained on the **Laval-Objaverse Dataset (LOD)**, comprising **90,545 high-quality 3D assets** and **39,008 diverse illumination conditions**.

- πŸ€— **Browse/Download the Dataset**: [vLAR/LavalObjaverseDataset](https://huggingface.co/datasets/vLAR/LavalObjaverseDataset)

- πŸ“– **Rendering Instructions**: See the [`RENDERING_INSTRUCTION.md`](https://github.com/vLAR-group/RelightFormer/blob/main/laval-objaverse-dataset/RENDERING_INSTRUCTION.md) in the GitHub repo.

---

## πŸ‹οΈ Training Details

RelightFormer is fine-tuned from the **Wan 2.1** base model. The training pipeline consists of two stages:

1. **Main Training**: Trained on the full LOD dataset to learn multi-view relighting priors.

1. **Post-Training**: A secondary fine-tuning stage to further enhance photorealism and consistency.

Detailed training configurations, hardware requirements (e.g., 4Γ— H200 GPUs), and scripts are available in the [GitHub Repository](https://github.com/vLAR-group/RelightFormer).

---

## πŸ“œ License

This model, its code, and associated datasets are licensed under the [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License](https://creativecommons.org/licenses/by-nc-sa/4.0/) (CC BY-NC-SA 4.0).

---

## πŸ™ Acknowledgements

This work was supported in part by the National Natural Science Foundation of China, the Research Grants Council of Hong Kong, the Otto Poon Charitable Foundation Smart Cities Research Institute, the Research Center for Unmanned Autonomous Systems, and the PolyU Kunpeng & Ascend Technology Innovation Incubation Center, The Hong Kong Polytechnic University.