File size: 5,831 Bytes
9e14838 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | # β Updated WACV 2026 paper announcement β
We are excited to announce that our [new paper](https://arxiv.org/abs/2508.06248) has been accepted to WACV 2026! The updated version includes additional experiments, models, and insights. Check out the latest version on [GitHub](https://github.com/yermandy/GenD).
---
## Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection
[](https://arxiv.org/abs/2503.19683)
[](https://huggingface.co/yermandy/deepfake-detection)
This is the official repository for the paper:
**[Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection](https://arxiv.org/abs/2503.19683)**.
### Abstract
> This paper tackles the challenge of detecting partially manipulated facial deepfakes, which involve subtle alterations to specific facial features while retaining the overall context, posing a greater detection difficulty than fully synthetic faces. We leverage the Contrastive Language-Image Pre-training (CLIP) model, specifically its ViT-L/14 visual encoder, to develop a generalizable detection method that performs robustly across diverse datasets and unknown forgery techniques with minimal modifications to the original model. The proposed approach utilizes parameter-efficient fine-tuning (PEFT) techniques, such as LN-tuning, to adjust a small subset of the model's parameters, preserving CLIP's pre-trained knowledge and reducing overfitting. A tailored preprocessing pipeline optimizes the method for facial images, while regularization strategies, including L2 normalization and metric learning on a hyperspherical manifold, enhance generalization. Trained on the FaceForensics++ dataset and evaluated in a cross-dataset fashion on Celeb-DF-v2, DFDC, FFIW, and others, the proposed method achieves competitive detection accuracy comparable to or outperforming much more complex state-of-the-art techniques. This work highlights the efficacy of CLIP's visual encoder in facial deepfake detection and establishes a simple, powerful baseline for future research, advancing the field of generalizable deepfake detection.
## Set up environment
``` bash
conda create --name dfdet python=3.12 uv
conda activate dfdet
uv pip install -r requirements.txt
```
## Minimal inference example
**β Important note**: sample images are already preprocessed. To get the same results as in the paper, you need to preprocess images using DeepfakeBench [preprocessing](https://github.com/SCLBD/DeepfakeBench/blob/fb6171a8e1db2ae0f017d9f3a12be31fd9e0a3fb/preprocessing/preprocess.py) pipeline.
### Minimal dependencies (torch + transformers)
This example requires only `torch` and `transformers` to run. This is an easy-to-integrate solution. The model has been traced and saved to a [`model.torchscript`](https://huggingface.co/yermandy/deepfake-detection/tree/main) file. Run:
``` bash
python inference_torchscript.py
```
Results might be a little bit different than in **precise inference** β
### Precise inference (full dependencies)
Read `inference.py`, it automatically downloads the model from [huggingface](https://huggingface.co/yermandy/deepfake-detection/tree/main) and runs inference on sample images.
``` bash
python inference.py
```
## Training
### Minimal example without external data
#### Run Training
You can adjust training configuration in `get_train_config` function in `run.py` or override them with command line arguments. Command line arguments have higher priority.
Example changing configurations in `get_train_config`:
1. Set `config.wandb = True` for logging to wandb
2. Set `config.devices = [2]` for using GPU number 2
``` bash
python run.py --train
```
#### Run testing (for example, on other dataset)
``` bash
python run.py --test
```
---
### Full training
#### Prepare the dataset
To fully train the model, you need to download datasets, preprocess them, and create a file with paths to the images.
For example, if you want to work with the [FaceForensics++](https://github.com/ondyari/FaceForensics) dataset, follow these steps:
1. Download the dataset first from the [official source](https://github.com/ondyari/FaceForensics)
2. Preprocess the dataset using [DeepfakeBench](https://github.com/SCLBD/DeepfakeBench)
3. Place images in the recommended directory structure: `datasets / <dataset_name> / <source_name> / <video_name> / <frame_name>`, see `src/dataset/deepfake.py` for more details
``` bash
datasets
βββ FF
βββ DF
β βββ 000_003
β βββ 025.png
β βββ 038.png
βββ F2F
β βββ 000_003
β βββ 019.png
β βββ 029.png
βββ FS
β βββ 000_003
β βββ 019.png
β βββ 029.png
βββ NT
β βββ 000_003
β βββ 019.png
β βββ 029.png
βββ real
βββ 000
βββ 025.png
βββ 038.png
```
4. Create files with paths to images similar to the ones in `config/datasets` directory. Get inspired by this script:
``` bash
sh scripts/prepare_FF.sh
```
#### Run training
Adjust training configuration as needed before executing the command below:
``` bash
python run.py --train
```
### Cite
``` bibtex
@article{yermakov-2025-deepfake-detection,
title={Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection},
author={Andrii Yermakov and Jan Cech and Jiri Matas},
year={2025},
eprint={2503.19683},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.19683},
}
```
|