Robust Overfitting Checkpoints: PreActResNet-18 on CIFAR-10

This repository contains PyTorch model checkpoints from our study investigating robust overfitting in PreActResNet-18 on CIFAR-10 across three multi-seed adversarial-training conditions: pixel-space PGD, low-frequency DCT-masked PGD, and mixed-domain training.

The checkpoints are intended for research, reproduction, and analysis of how clean and adversarial robustness change throughout training.

Repository structure

KaiwenDu/robust-overfitting-checkpoints/
β”œβ”€β”€ pixel-only/                       # Pixel-space PGD-10 multi-seed runs (40 checkpoints each)
β”‚   β”œβ”€β”€ seed-42/
β”‚   β”œβ”€β”€ seed-43/
β”‚   β”œβ”€β”€ seed-44/
β”‚   β”œβ”€β”€ seed-45/
β”‚   └── seed-46/
β”œβ”€β”€ low-frequency-only/               # Low-frequency DCT-masked PGD-10 (Cutoff = 8, 40 checkpoints each)
β”‚   β”œβ”€β”€ seed-42/
β”‚   β”œβ”€β”€ seed-43/
β”‚   β”œβ”€β”€ seed-44/
β”‚   β”œβ”€β”€ seed-45/
β”‚   └── seed-46/
└── mixed-domain/                     # Mixed pixel/DCT PGD-10 multi-seed runs (40 checkpoints each)
    β”œβ”€β”€ seed-42/
    β”œβ”€β”€ seed-43/
    β”œβ”€β”€ seed-44/
    β”œβ”€β”€ seed-45/
    └── seed-46/

Model details

  • Architecture: PreActResNet-18
  • Task: CIFAR-10 image classification
  • Classes: 10
  • Training conditions:
    • Pixel-only: Pixel-space PGD-10 across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in pixel-only/seed-<seed>/)
    • Low-frequency only: DCT-masked PGD-10 ($k=8$ cutoff) across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in low-frequency-only/seed-<seed>/)
    • Mixed-domain: Seeded 50/50 epoch-level randomized selection between pixel-space PGD-10 and low-frequency DCT-masked PGD-10 ($k=8$) across 5 seeds: 42, 43, 44, 45, 46 (40 checkpoints each in mixed-domain/seed-<seed>/)
  • Training duration: 200 epochs per run
  • Checkpoints: Every 5 epochs (epoch_5.pt through epoch_200.pt)
  • Perturbation budget: (L_\infty) epsilon = 8/255
  • PGD step size: 2/255

The model architecture and full training code are available in the companion GitHub repository:
https://github.com/ItsKaiwenDu/Robust-Overfitting

Results

Each checkpoint was evaluated on the full CIFAR-10 test set using 20-step PGD (both standard pixel-space PGD and DCT-masked low-frequency PGD with cutoff $k=8$) with epsilon = 8/255 and step size = 2/255. Joint robustness measures the percentage of test samples classified correctly under both evaluated attack types.

Evaluation Condition / Checkpoint Clean accuracy Pixel-PGD-20 robust accuracy Low-Freq PGD-20 robust accuracy Joint robustness
Pixel-Only Multi-Seed (Seeds 42–46 Mean Β± SD)
Epoch 105 (best pixel robustness) 83.04% Β± 0.33% 51.22% Β± 0.28% 77.70% Β± 0.27% 51.22% Β± 0.28%
Epoch 200 (final checkpoint) 84.40% Β± 0.09% 42.66% Β± 0.22% 76.48% Β± 0.26% 42.66% Β± 0.22%
Low-Frequency-Only Multi-Seed (Seeds 42–46 Mean Β± SD)
Epoch 195 (best low-frequency robustness) 94.24% Β± 0.19% 0.00% Β± 0.00% 92.74% Β± 0.17% 0.00% Β± 0.00%
Epoch 200 (final checkpoint) 94.28% Β± 0.22% 0.00% Β± 0.00% 92.70% Β± 0.28% 0.00% Β± 0.00%
Mixed-Domain Multi-Seed (Seeds 42–46 Mean Β± SD)
Epoch 85 (best pixel & joint robustness) 71.86% Β± 0.88% 40.80% Β± 0.78% 66.16% Β± 0.80% 40.79% Β± 0.79%
Epoch 200 (final checkpoint) 90.33% Β± 3.45% 19.33% Β± 14.28% 85.20% Β± 4.29% 19.33% Β± 14.27%

Key Observations

  • Robust Overfitting in Pixel Training: All five pixel-only runs peak under Pixel-PGD-20 at epoch 105 (51.22% Β± 0.28%) and decline by 8.56 percentage points by epoch 200, while clean accuracy improves to 84.40%.
  • Low-Frequency Training Dynamics: Low-frequency-only training yields strong matched-domain robustness (92.74% at epoch 195; 92.70% at epoch 200) with no practically meaningful late decline. It does not transfer to unrestricted pixel perturbations (0.00%).
  • Mixed-Domain Trade-offs and Dynamics: Alternating between pixel-space and low-frequency attacks does not yield stable joint robustness. Pixel robustness is strongly associated with the immediately preceding training domain (point-biserial $r = 0.971$), so the epoch-85-to-200 change is schedule-confounded and should not be interpreted as pure robust overfitting.

Which checkpoint should I use?

  • For Pixel-Space Robustness (Pixel-Only): Use epoch_105.pt for peak pixel and joint robustness, or epoch_200.pt to evaluate the final model after robust overfitting.
  • For Low-Frequency Defense (Low-Frequency-Only): Use epoch_195.pt for the highest measured low-frequency robustness (92.74%), or epoch_200.pt for the highest measured clean accuracy (94.28%).
  • For Mixed-Domain Robustness (Mixed-Domain): epoch_85.pt records the highest aggregate pixel and joint robustness, but it follows a pixel-training epoch. The checkpoints do not provide stable simultaneous robustness to both threats.
  • For Dynamics Research: Use the full 40-checkpoint sequence (epoch_5.pt to epoch_200.pt) to reproduce and analyze clean and robust accuracy curves across training.

Loading a checkpoint

Clone the companion code repository first, since it contains the PreActResNet-18 definition:

git clone https://github.com/ItsKaiwenDu/Robust-Overfitting.git
cd Robust-Overfitting
pip install -r requirements.txt

Then download and load a checkpoint:

import torch
from huggingface_hub import hf_hub_download
from models.preact_resnet import PreActResNet18

# Example: Download Epoch 85 from Mixed-Domain training (Seed 42)
checkpoint_path = hf_hub_download(
    repo_id="KaiwenDu/robust-overfitting-checkpoints",
    filename="epoch_85.pt",
    subfolder="mixed-domain/seed-42",
)

# Example: Download Epoch 105 from Pixel-Only training (Seed 42)
# checkpoint_path = hf_hub_download(
#     repo_id="KaiwenDu/robust-overfitting-checkpoints",
#     filename="epoch_105.pt",
#     subfolder="pixel-only/seed-42",
# )

# Example: Download Epoch 105 from Low-Frequency training (Seed 42)
# checkpoint_path = hf_hub_download(
#     repo_id="KaiwenDu/robust-overfitting-checkpoints",
#     filename="epoch_105.pt",
#     subfolder="low-frequency-only/seed-42",
# )

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = PreActResNet18(num_classes=10).to(device)

checkpoint = torch.load(checkpoint_path, map_location=device)
model.load_state_dict(checkpoint["model_state_dict"])
model.eval()

Inputs should be CIFAR-10 RGB images converted to tensors in [0, 1] and normalized with:

mean = (0.4914, 0.4822, 0.4465)
std = (0.2471, 0.2435, 0.2616)

Limitations

These checkpoints represent empirical research studies on robust overfitting, not a claim of state-of-the-art adversarial robustness. Robustness was measured against specified pixel-space and DCT-masked low-frequency PGD-20 attacks; it should not be interpreted as robustness against every possible attack.

Citation

If you use these checkpoints, please cite the companion repository and the original robust-overfitting paper:

@article{rice2020overfitting,
  title={Overfitting in Adversarially Robust Deep Learning},
  author={Rice, Leslie and Wong, Eric and Kolter, J. Zico},
  journal={Proceedings of the 37th International Conference on Machine Learning},
  year={2020}
}

License

The companion code is released under the MIT License. CIFAR-10 is subject to its own dataset terms and license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train KaiwenDu/robust-overfitting-checkpoints