Robust Overfitting Checkpoints: PreActResNet-18 on CIFAR-10
This repository contains PyTorch model checkpoints from our study investigating robust overfitting in PreActResNet-18 on CIFAR-10 across three multi-seed adversarial-training conditions: pixel-space PGD, low-frequency DCT-masked PGD, and mixed-domain training.
The checkpoints are intended for research, reproduction, and analysis of how clean and adversarial robustness change throughout training.
Repository structure
KaiwenDu/robust-overfitting-checkpoints/
βββ pixel-only/ # Pixel-space PGD-10 multi-seed runs (40 checkpoints each)
β βββ seed-42/
β βββ seed-43/
β βββ seed-44/
β βββ seed-45/
β βββ seed-46/
βββ low-frequency-only/ # Low-frequency DCT-masked PGD-10 (Cutoff = 8, 40 checkpoints each)
β βββ seed-42/
β βββ seed-43/
β βββ seed-44/
β βββ seed-45/
β βββ seed-46/
βββ mixed-domain/ # Mixed pixel/DCT PGD-10 multi-seed runs (40 checkpoints each)
βββ seed-42/
βββ seed-43/
βββ seed-44/
βββ seed-45/
βββ seed-46/
Model details
- Architecture: PreActResNet-18
- Task: CIFAR-10 image classification
- Classes: 10
- Training conditions:
- Pixel-only: Pixel-space PGD-10 across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in
pixel-only/seed-<seed>/) - Low-frequency only: DCT-masked PGD-10 ($k=8$ cutoff) across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in
low-frequency-only/seed-<seed>/) - Mixed-domain: Seeded 50/50 epoch-level randomized selection between pixel-space PGD-10 and low-frequency DCT-masked PGD-10 ($k=8$) across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in
mixed-domain/seed-<seed>/)
- Pixel-only: Pixel-space PGD-10 across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in
- Training duration: 200 epochs per run
- Checkpoints: Every 5 epochs (
epoch_5.ptthroughepoch_200.pt) - Perturbation budget: (L_\infty) epsilon = 8/255
- PGD step size: 2/255
The model architecture and full training code are available in the companion GitHub repository:
https://github.com/ItsKaiwenDu/Robust-Overfitting
Results
Each checkpoint was evaluated on the full CIFAR-10 test set using 20-step PGD (both standard pixel-space PGD and DCT-masked low-frequency PGD with cutoff $k=8$) with epsilon = 8/255 and step size = 2/255. Joint robustness measures the percentage of test samples classified correctly under both evaluated attack types.
| Evaluation Condition / Checkpoint | Clean accuracy | Pixel-PGD-20 robust accuracy | Low-Freq PGD-20 robust accuracy | Joint robustness |
|---|---|---|---|---|
| Pixel-Only Multi-Seed (Seeds 42β46 Mean Β± SD) | ||||
| Epoch 105 (best pixel robustness) | 83.04% Β± 0.33% | 51.22% Β± 0.28% | 77.70% Β± 0.27% | 51.22% Β± 0.28% |
| Epoch 200 (final checkpoint) | 84.40% Β± 0.09% | 42.66% Β± 0.22% | 76.48% Β± 0.26% | 42.66% Β± 0.22% |
| Low-Frequency-Only Multi-Seed (Seeds 42β46 Mean Β± SD) | ||||
| Epoch 195 (best low-frequency robustness) | 94.24% Β± 0.19% | 0.00% Β± 0.00% | 92.74% Β± 0.17% | 0.00% Β± 0.00% |
| Epoch 200 (final checkpoint) | 94.28% Β± 0.22% | 0.00% Β± 0.00% | 92.70% Β± 0.28% | 0.00% Β± 0.00% |
| Mixed-Domain Multi-Seed (Seeds 42β46 Mean Β± SD) | ||||
| Epoch 85 (best pixel & joint robustness) | 71.86% Β± 0.88% | 40.80% Β± 0.78% | 66.16% Β± 0.80% | 40.79% Β± 0.79% |
| Epoch 200 (final checkpoint) | 90.33% Β± 3.45% | 19.33% Β± 14.28% | 85.20% Β± 4.29% | 19.33% Β± 14.27% |
Key Observations
- Robust Overfitting in Pixel Training: All five pixel-only runs peak under Pixel-PGD-20 at epoch 105 (51.22% Β± 0.28%) and decline by 8.56 percentage points by epoch 200, while clean accuracy improves to 84.40%.
- Low-Frequency Training Dynamics: Low-frequency-only training yields strong matched-domain robustness (92.74% at epoch 195; 92.70% at epoch 200) with no practically meaningful late decline. It does not transfer to unrestricted pixel perturbations (0.00%).
- Mixed-Domain Trade-offs and Dynamics: Alternating between pixel-space and low-frequency attacks does not yield stable joint robustness. Pixel robustness is strongly associated with the immediately preceding training domain (point-biserial $r = 0.971$), so the epoch-85-to-200 change is schedule-confounded and should not be interpreted as pure robust overfitting.
Which checkpoint should I use?
- For Pixel-Space Robustness (Pixel-Only): Use
epoch_105.ptfor peak pixel and joint robustness, orepoch_200.ptto evaluate the final model after robust overfitting. - For Low-Frequency Defense (Low-Frequency-Only): Use
epoch_195.ptfor the highest measured low-frequency robustness (92.74%), orepoch_200.ptfor the highest measured clean accuracy (94.28%). - For Mixed-Domain Robustness (Mixed-Domain):
epoch_85.ptrecords the highest aggregate pixel and joint robustness, but it follows a pixel-training epoch. The checkpoints do not provide stable simultaneous robustness to both threats. - For Dynamics Research: Use the full 40-checkpoint sequence (
epoch_5.pttoepoch_200.pt) to reproduce and analyze clean and robust accuracy curves across training.
Loading a checkpoint
Clone the companion code repository first, since it contains the PreActResNet-18 definition:
git clone https://github.com/ItsKaiwenDu/Robust-Overfitting.git
cd Robust-Overfitting
pip install -r requirements.txt
Then download and load a checkpoint:
import torch
from huggingface_hub import hf_hub_download
from models.preact_resnet import PreActResNet18
# Example: Download Epoch 85 from Mixed-Domain training (Seed 42)
checkpoint_path = hf_hub_download(
repo_id="KaiwenDu/robust-overfitting-checkpoints",
filename="epoch_85.pt",
subfolder="mixed-domain/seed-42",
)
# Example: Download Epoch 105 from Pixel-Only training (Seed 42)
# checkpoint_path = hf_hub_download(
# repo_id="KaiwenDu/robust-overfitting-checkpoints",
# filename="epoch_105.pt",
# subfolder="pixel-only/seed-42",
# )
# Example: Download Epoch 105 from Low-Frequency training (Seed 42)
# checkpoint_path = hf_hub_download(
# repo_id="KaiwenDu/robust-overfitting-checkpoints",
# filename="epoch_105.pt",
# subfolder="low-frequency-only/seed-42",
# )
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = PreActResNet18(num_classes=10).to(device)
checkpoint = torch.load(checkpoint_path, map_location=device)
model.load_state_dict(checkpoint["model_state_dict"])
model.eval()
Inputs should be CIFAR-10 RGB images converted to tensors in [0, 1] and normalized with:
mean = (0.4914, 0.4822, 0.4465)
std = (0.2471, 0.2435, 0.2616)
Limitations
These checkpoints represent empirical research studies on robust overfitting, not a claim of state-of-the-art adversarial robustness. Robustness was measured against specified pixel-space and DCT-masked low-frequency PGD-20 attacks; it should not be interpreted as robustness against every possible attack.
Citation
If you use these checkpoints, please cite the companion repository and the original robust-overfitting paper:
@article{rice2020overfitting,
title={Overfitting in Adversarially Robust Deep Learning},
author={Rice, Leslie and Wong, Eric and Kolter, J. Zico},
journal={Proceedings of the 37th International Conference on Machine Learning},
year={2020}
}
License
The companion code is released under the MIT License. CIFAR-10 is subject to its own dataset terms and license.