File size: 12,282 Bytes
9e14838
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
<p align="center">
  <h1 align="center">[CVPR'25] Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Model (DFD-FCG)</h1>

  <p align="center">
    <a href="https://github.com/ODD2"><strong>Yue-Hua Han</strong></a>
    <sup>1,3,4</sup>
    &nbsp;&nbsp;
    <a href="https://github.com/Teddy12155555"><strong>Tai-Ming Huang</strong></a>
    <sup>1,3,4</sup>
    &nbsp;&nbsp;
    <a href="https://scholar.google.com/citations?user=nnzQtDAAAAAJ&hl=zh-TW"><strong>Kai-Lung Hua</strong></a>
    <sup>2,4</sup>
    &nbsp;&nbsp;
    <a href="https://scholar.google.com.au/citations?user=3x9KITUAAAAJ&hl=en"><strong> Jun-Cheng Chen</strong></a>
    <sup>1</sup>
    <br>
    <!-- <sup>1</sup>Academia Sinica&nbsp;
    <sup>2</sup>Microsoft&nbsp;
    <sup>3</sup>National Taiwan University&nbsp;
    <br>
    <sup>4</sup>National Taiwan University of Science and Technology&nbsp; -->
    <sup>1</sup>Academia Sinica,</span>&nbsp;
    <sup>2</sup>Microsoft,</span>&nbsp;
    <sup>3</sup>National Taiwan University,</span>&nbsp;
    <br>
    <sup>4</sup>National Taiwan University of Science and Technology</span>&nbsp;
    <br>
    <a href='https://arxiv.org/abs/2404.05583'><img src='https://img.shields.io/badge/ArXiv-2404.05583-red'></a>&nbsp;
    <img src="assets/teaser.png">
  </p>
</p>

## πŸ₯‡Abstract
<div style="text-align: justify">  
Generative models have enabled the creation of highly realistic facial-synthetic images, raising significant concerns due to their potential for misuse. Despite rapid advancements in the field of deepfake detection, developing efficient approaches to leverage foundation models for improved generalizability to unseen forgery samples remains challenging. To address this challenge, we propose a novel side-network-based decoder that extracts spatial and temporal cues using the CLIP image encoder for generalized video-based Deepfake detection. Additionally, we introduce Facial Component Guidance (FCG) to enhance spatial learning generalizability by encouraging the model to focus on key facial regions. By leveraging the generic features of a vision-language foundation model, our approach demonstrates promising generalizability on challenging Deepfake datasets while also exhibiting superiority in training data efficiency, parameter efficiency, and model robustness.
</div>

## πŸ“TODOs
  - [x] Training + Evaluation Code
  - [x] Model Weights
  - [x] Inference Code
  - [ ] HeyGen Evaluation Dataset


## πŸ™ŒNews
  - June 08: We have released the [model checkpoint](https://drive.google.com/file/d/1ydD5rnaaF0i2zLE7NidLtAhjonHoVQOk/view?usp=sharing) and provided inference code for single videos! Checkout [this section](#inference---demo-video) for further details!

## πŸš€Installation
```shell
# conda environment
conda env create -f environment.yml
```

## πŸ“‚Dataset Structure
The structure of the **pre-processed datasets** for our project, the video files (*.avi) have been processed to only retain the aligned face. We use soft-links **(ln -s)** to manage and link the folders containing pre-processed videos on different drives.
```shell
datasets
β”œβ”€β”€ cdf
β”‚   β”œβ”€β”€ FAKE
β”‚   β”‚   └── videos
β”‚   β”‚         └── *.avi
β”‚   β”œβ”€β”€ REAL
β”‚   β”‚   └── videos
β”‚   β”‚         └── *.avi
β”‚   └── csv_files
β”‚       β”œβ”€β”€ test_fake.csv
β”‚       └── test_real.csv
β”œβ”€β”€ dfdc
β”‚   β”œβ”€β”€ csv_files
β”‚   β”‚   └── test.csv
β”‚   └── videos
β”œβ”€β”€ dfo
β”‚   β”œβ”€β”€ FAKE
β”‚   β”‚   └── videos
β”‚   β”‚         └── *.avi
β”‚   β”œβ”€β”€ REAL
β”‚   β”‚   └── videos
β”‚   β”‚         └── *.avi
β”‚   └── csv_files
β”‚       β”œβ”€β”€ test_fake.csv
β”‚       └── test_real.csv
β”œβ”€β”€ ffpp
β”‚   β”œβ”€β”€ DF
β”‚   β”‚   β”œβ”€β”€ c23
β”‚   β”‚   β”‚   └── videos
β”‚   β”‚   β”‚         └── *.avi
β”‚   β”‚   β”œβ”€β”€ c40
β”‚   β”‚   β”‚   └── videos
β”‚   β”‚   β”‚         └── *.avi
β”‚   β”‚   └── raw
β”‚   β”‚       └── videos
β”‚   β”‚   β”‚         └── *.avi
β”‚   β”œβ”€β”€ F2F ...
β”‚   β”œβ”€β”€ FS ...
β”‚   β”œβ”€β”€ FSh ...
β”‚   β”œβ”€β”€ NT ...
β”‚   β”œβ”€β”€ real ...
β”‚   └── csv_files
β”‚       β”œβ”€β”€ test.json
β”‚       β”œβ”€β”€ train.json
β”‚       └── val.json
|   
└── robustness
    β”œβ”€β”€ BW
    β”‚   β”œβ”€β”€ 1
    β”‚   β”‚   β”œβ”€β”€ DF
    β”‚   β”‚   β”‚   └── c23
    β”‚   β”‚   β”‚       └── videos
    β”‚   β”‚   β”‚             └── *.avi
    β”‚   β”‚   β”œβ”€β”€ F2F ...
    β”‚   β”‚   β”œβ”€β”€ FS ...
    β”‚   β”‚   β”œβ”€β”€ FSh ...
    β”‚   β”‚   β”œβ”€β”€ NT ...
    β”‚   β”‚   β”œβ”€β”€ real ...
    β”‚   β”‚   β”‚
    β”‚   β”‚   └── csv_files
    β”‚   β”‚       β”œβ”€β”€ test.json
    β”‚   β”‚       β”œβ”€β”€ train.json
    β”‚   β”‚       └── val.json
    β”‚   β”‚   
    β”‚   β”‚   
    β”‚   β”‚   
    .   .
    .   .
    .   .
```


## πŸ”§Dataset Pre-processing
### Generic Pre-processing
This phase performs the required pre-processing for our method, this includes *facial alignment (using the mean face from LRW)* and *facial cropping*.
```bash
# First, fetch all the landmarks & bboxes of the video frames.
python -m src.preprocess.fetch_landmark_bbox \ 
--root-dir="/storage/FaceForensicC23" \ # The root folder of the dataset
--video-dir="videos" \  # The root folder of the videos
--fdata-dir="frame_data" \ # The folder to save the extracted frame data
--glob-exp="*/*" \  # The glob expression to search through the root video folder
--split-num=1 \ # Split the dataset into several parts for parallel process.
--part-num=1 \ # The part of dataset to process  for parallel process.
--batch=1 \ # The batch size for the 2D-FAN face data extraction. (suggestion: 1)
--max-res=800  # The maximum resolution for either side of the image

# Then, crop all the faces from the original videos.
python -m src.preprocess.crop_main_face \ 
--root-dir="/storage/FaceForensicC23/" \ # The root folder of the dataset
--video-dir="videos" \  # The root folder of the videos
--fdata-dir="frame_data" \ # The folder to fetch the frame data for landmarks and bboxes
--glob-exp="*/*" \  # The glob expression to search through the root video folder
--crop-dir="cropped" \ # The folder to save the cropped videos
--crop-width=150 \ # The width for the cropped videos
--crop-height=150 \ # The height for the cropped videos
--mean-face="./misc/20words_mean_face.npy" # The mean face for face aligned cropping. 
--replace \ # Control whether to replace existing cropped videos
--workers=1 # Number of works to perform parallel process (default: cpu / 2 )
```


### Robustness Pre-processing
This phase requires the pre-processed facial landmarks to perform facial cropping, please refer to the **Generic Pre-processing** for further detail.
```bash
# First, we add perturbation to all the videos.
python -m src.preprocess.phase1_apply_all_to_videos \ 
--dts-root="/storage/FaceForensicC23" \ # The root folder of the dataset
--vid-dir="videos" \  # The root folder of the videos
--rob-dir="robustness" \ # The folder to save the perturbed videos
--glob-exp="*/*.mp4" \  # The glob expression to search through the root video folder
--split=1 \ # Split the dataset into several parts for parallel process.
--part=1 \ # The part of dataset to process  for parallel process.
--workers=1  # Number of works to perform parallel process (default: cpu / 2 )

# Then, crop all the faces from the perturbed videos.
python -m src.preprocess.phase2_face_crop_all_videos \ 
(setup/run/clean) # the three phase operations
--root-dir="/storage/FaceForensicC23/" \ # The root folder of the dataset
--rob-dir="videos" \  # The root folder of the robustness videos
--fd-dir="frame_data" \ # The folder to fetch the frame data for landmarks and bboxes
--glob-exp="*/*/*/*.mp4" \  # The glob expression to search through the root video folder
--crop-dir="cropped_robust" \ # The folder to save the cropped videos
--mean-face="./misc/20words_mean_face.npy" \ # The mean face for face aligned cropping. 
--workers=1  # Number of works to perform parallel process (default: cpu / 2 )
```

## πŸ€–Training & Evaluation
### Training - Preset Settings 
In `./scripts`,  scripts are provided to start the training process for the settings mentioned in our paper. 
These settings are configured to run on a cluster with `V100*4`.
```bash
bash ./scripts/model/ffg_l14.sh # begin training process
```
### Training - Custom Settings
Our project is built on `pytorch-lightning (2.2.0)`, please refer the  [official manual](https://lightning.ai/docs/pytorch/2.2.0/common/trainer.html#trainer-class-api) and adjust the following files for advance configurations:
```bash
./configs/base.yaml # major training settings (e.g. epochs, optimizer, batch size, mixed-precision ...)
./configs/data.yaml # settings for the training & validation dataset
./configs/inference.yaml # settings for the evaluation dataset (extension of data.yaml)
./configs/logger.yaml # settings for the WandB logger
./configs/clip/L14/ffg.yaml # settings for the main model 
./configs/test.yaml # settings for debugging (offline logging, small batch size, short epochs ...)
```
The following command starts the training process with the provided settings:
```bash
# For debugging, add '--config configs/test.yaml' after the '--config configs/clip/L14/ffg.yaml'.
python main.py \
--config configs/base.yaml \
--config configs/clip/L14/ffg.yaml

# Fine-grained control is supported with the pytorch-lightning-cli.
python main.py \
--config configs/base.yaml \
--config configs/clip/L14/ffg.yaml \
--optimizer.lr=1e-5 \
--trainer.max_epochs=10 \
--data.init_args.train_datamodules.init_args.batch_size=5
```
### Evaluation - Standard
To perform evaluation on datasets, run the following command:
```bash
python inference.py \
"logs/fcg_l14/setting.yaml" \ # model settings
"./configs/inference.yaml" \ # evaluation dataset settings
"logs/fcg_l14/checkpoint.ckpt" \ # model checkpoint
"--devices=4" # number of devices to compute in parallel
```
### Evaluation - Robustness
We provide tools in `./scripts/tools/` to simplify the robustness evaluation task: `create-robust-config.sh` creates an evaluation config for each perturbation types and `inference-robust.sh` runs through all the datasets with the specified model.

## 😎Inference - Demo Video
To run inference on a single video with an indicator, please download our model checkpoint and execute the following commands:
```bash
# Pre-Processing: fetch facial landmark and bounding box
python -m src.preprocess.fetch_landmark_bbox \
--root-dir="./resources" \
--video-dir="videos" \
--fdata-dir="frame_data" \
--glob-exp="*"
# Pre-Processing: crop out the facial regions
python -m src.preprocess.crop_main_face \
--root-dir="./resources" \
--video-dir="videos" \
--fdata-dir="frame_data" \
--crop-dir="cropped" \
--glob-exp="*"
# Main Process
python -m demo \
"checkpoint/setting.yaml" \ # the model setting of the checkpoint
"checkpoint/weights.ckpt" \ # the model weights of the checkpoint
"resources/videos/000_003.mp4" \ # the video to process
--out_path="test.avi" \ # the output path of the processed video
--threshold=0.5 \ # the threshold for the real/fake indicator
--batch_size=30 # the input batch size of the model (~10G VRAM when batch_size=30 )
```
The following is a sample frame from the processed video:
<p align="center">
    <img src="assets/demo.png">
</p>


<!-- ## πŸ”₯ Inference -->
## πŸ”— BibTeX
If you find our efforts helpful, please cite our paper and leave a star for further updates!
```bibtex

@inproceedings{cvpr25_dfd_fcg,
      title={Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Model},
      author={Yue-Hua Han, Tai-Ming Huang, Kai-Lung Hua, Jun-Cheng Chen},
      booktitle={Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2025}
}
```


## πŸ“­ Contact
The provided code and weights are only available for research purpose only. 
If you have further questions (including commercial use), please contact [Dr. Jun-Cheng Chen](pullpull@citi.sinica.edu.tw).