File size: 5,273 Bytes
7163597
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
---
license: apache-2.0
pipeline_tag: image-classification
library_name: pytorch
base_model: openai/clip-vit-large-patch14
tags:
- ai-generated-image-detection
- deepfake-detection
- synthetic-image-detection
- image-forensics
- clip
- vision-language-model
- prompt-learning
- pytorch-lightning
- arxiv:2408.08412
---

# PoundNet

PoundNet checkpoint weights for the paper **"Penny-Wise and Pound-Foolish in AI-Generated Image Detection"**.

PoundNet is a CLIP-based AI-generated image detector built around asymmetric prompt learning for binary real/fake classification and category-aware supervision. The method is designed to reduce the "penny-wise and pound-foolish" behavior of deepfake detectors: strong performance on a narrow training distribution but poor generalization and degraded upstream semantic knowledge.

These weights are released for use with the official PoundNet codebase:

- Code: https://github.com/iamwangyabin/PoundNet
- arXiv: https://arxiv.org/abs/2408.08412
- Model weights: https://huggingface.co/nebula/PoundNet

## Model Details

- **Architecture**: PoundNet
- **Backbone**: CLIP ViT-L/14
- **Task**: binary AI-generated image detection / deepfake detection
- **Output**: real/fake prediction scores through the official evaluation code
- **Training data**: ProGAN split from the ForenSynths-style training setup used by the official PoundNet implementation
- **Checkpoint format**: PyTorch Lightning `.ckpt`

The checkpoints in this repository are not standalone `transformers` checkpoints. They should be loaded with the official PoundNet repository and configuration files.

## Released Checkpoints

| Checkpoint | File |
| --- | --- |
| `poundnet_ViTL_Progan_20240506_23_30_25` | `poundnet_ViTL_Progan_20240506_23_30_25/last.ckpt` |
| `poundnet_ViTL_Progan_20240804_21_16_47` | `poundnet_ViTL_Progan_20240804_21_16_47/last.ckpt` |
| `poundnet_ViTL_Progan_20240805_10_31_08` | `poundnet_ViTL_Progan_20240805_10_31_08/last.ckpt` |
| `poundnet_ViTL_Progan_20240805_12_09_21` | `poundnet_ViTL_Progan_20240805_12_09_21/last.ckpt` |

## Installation

Clone the official repository and install dependencies:

```bash
git clone https://github.com/iamwangyabin/PoundNet.git
cd PoundNet
pip install -r requirements.txt
```

Install PyTorch separately according to your CUDA environment before installing the remaining dependencies.

## Download Weights

```bash
mkdir -p weights

wget -O ./weights/poundnet_ViTL_Progan_20240506_23_30_25.ckpt \
  https://huggingface.co/nebula/PoundNet/resolve/main/poundnet_ViTL_Progan_20240506_23_30_25/last.ckpt

wget -O ./weights/poundnet_ViTL_Progan_20240804_21_16_47.ckpt \
  https://huggingface.co/nebula/PoundNet/resolve/main/poundnet_ViTL_Progan_20240804_21_16_47/last.ckpt

wget -O ./weights/poundnet_ViTL_Progan_20240805_10_31_08.ckpt \
  https://huggingface.co/nebula/PoundNet/resolve/main/poundnet_ViTL_Progan_20240805_10_31_08/last.ckpt

wget -O ./weights/poundnet_ViTL_Progan_20240805_12_09_21.ckpt \
  https://huggingface.co/nebula/PoundNet/resolve/main/poundnet_ViTL_Progan_20240805_12_09_21/last.ckpt
```

## Evaluation

PoundNet expects benchmark datasets saved in Hugging Face Arrow format and loaded through `datasets.load_from_disk(...)`. See the official repository for the expected dataset layout and download helper.

Example evaluation command:

```bash
python test.py --cfg cfgs/poundnet.yaml \
  datasets.base_path=/path/to/DF-arrow
```

The default evaluation config uses the ViT-L/14 PoundNet checkpoint and evaluates on multiple AI-generated image detection benchmarks through the official codebase.

## Intended Uses

PoundNet is intended for academic research on:

- AI-generated image detection
- deepfake detection
- synthetic image forensics
- cross-generator and cross-dataset generalization
- prompt-learning adaptation of vision-language models

## Limitations

These checkpoints are research artifacts and should not be treated as a complete production moderation or forensic system. Performance can vary under distribution shifts such as unseen generators, image editing pipelines, social media compression, resizing, screenshots, adversarial post-processing, or domain-specific content.

The released checkpoints require the official PoundNet code and configuration files. They are not directly loadable through `AutoModel.from_pretrained`.

## Ethical Considerations

PoundNet is released to support research on synthetic media detection and trustworthy image forensics. Users should validate performance carefully before applying it to real-world moderation, legal, journalistic, or security workflows. Detection results should not be used as the sole evidence for high-stakes decisions.

## Citation

If you use PoundNet, please cite:

```bibtex
@article{wang2026pennywise,
  title   = {Penny-Wise and Pound-Foolish in AI-Generated Image Detection},
  author  = {Wang, Yabin and Huang, Zhiwu and Su, Zhou and Prugel-Bennett, Adam and Hong, Xiaopeng},
  journal = {IEEE Transactions on Pattern Analysis and Machine Intelligence},
  pages   = {1--14},
  year    = {2026},
  doi     = {10.1109/TPAMI.2026.3664388}
}
```

## Links

- Code: https://github.com/iamwangyabin/PoundNet
- arXiv: https://arxiv.org/abs/2408.08412
- Weights: https://huggingface.co/nebula/PoundNet