RVL-CDIP document classifiers, with and without identification codes
Seven image classifiers and three linear SVM baselines trained on RVL-CDIP (16 document categories), each in four versions that differ in the labels and in the images they were trained on:
| folder | labels | training images |
|---|---|---|
labels-original_codes-kept |
original RVL-CDIP labels | de-identified pages |
labels-original_codes-removed |
original RVL-CDIP labels | de-identified pages with the identification codes whitened out |
labels-corrected_codes-kept |
corrected labels | de-identified pages |
labels-corrected_codes-removed |
corrected labels | de-identified pages with the identification codes whitened out |
Identification codes are the Bates numbers stamped on most pages; they are associated with the
category label and act as a shortcut (see stefan-hf/rvlcdip-id-codes).
The codes were located with stefan-hf/yolov8n-rvlcdip-idcodes.
All models were trained on a de-identified copy of the corpus, in which personal data had been
replaced with synthetic values, so none of them was trained on the original personal data.
The codes-removed models cannot use the codes, at a cost of 0.0β0.3 points in-domain on the
corrected labels. On RVL-CDIP-N, which was collected from other sources, the codes-kept models
are 1β6 points more accurate (second table below), so neither version is better on every measure.
Models and accuracy
Test accuracy (%) of the released checkpoint on the test split of its own training condition; in parentheses, the mean over three training seeds. The released checkpoint is always seed 42, the first run, chosen in advance and not by score. Models trained on the original labels are scored on the original test labels (39,999 pages) and models trained on the corrected labels on the corrected test labels (36,446 pages), so the two halves of the table are not directly comparable.
| model | folder | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|
| AlexNet | alexnet |
89.21 (89.30) | 88.41 (88.52) | 92.85 (92.82) | 92.73 (92.73) | 228 MB |
| GoogLeNet | googlenet |
89.34 (89.25) | 88.52 (88.50) | 93.19 (93.15) | 92.91 (92.97) | 23 MB |
| ResNet-50 | resnet50 |
91.01 (90.85) | 90.21 (90.20) | 94.26 (94.31) | 94.28 (94.19) | 94 MB |
| ResNeXt-50 (32x4d) | resnext50 |
91.48 (91.33) | 90.48 (90.63) | 94.64 (94.69) | 94.56 (94.52) | 92 MB |
| SqueezeNet 1.0 | squeezenet |
88.70 (88.66) | 87.74 (87.69) | 92.46 (92.59) | 92.50 (92.46) | 3 MB |
| VGG-16 | vgg16 |
91.18 (91.27) | 90.74 (90.81) | 94.61 (94.62) | 94.46 (94.42) | 537 MB |
| DiT-base | dit_base |
92.69 (92.76) | 92.20 (92.19) | 95.94 (95.92) | 95.74 (95.81) | 343 MB |
Corrected-label models on both test conditions and on RVL-CDIP-N (1,002 newly collected pages):
| model | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed |
|---|---|---|---|---|---|---|
| AlexNet | 92.85 | 91.54 | 92.37 | 92.73 | 73.6 | 69.7 |
| GoogLeNet | 93.19 | 92.62 | 92.72 | 92.91 | 77.9 | 71.5 |
| ResNet-50 | 94.26 | 93.71 | 94.19 | 94.28 | 79.2 | 77.0 |
| ResNeXt-50 (32x4d) | 94.64 | 94.10 | 94.45 | 94.56 | 80.8 | 76.2 |
| SqueezeNet 1.0 | 92.46 | 91.57 | 92.32 | 92.50 | 77.0 | 75.0 |
| VGG-16 | 94.61 | 93.75 | 94.25 | 94.46 | 78.5 | 77.8 |
| DiT-base | 95.94 | 95.37 | 95.71 | 95.74 | 87.1 | 83.9 |
"kept β removed" is a model trained with codes and tested on pages without them.
Usage
rvlcdip_models.py in this repository builds the architecture, loads the weights, and applies the
training-time preprocessing (grayscale, bilinear resize to 224 Γ 224, replicate to three channels,
normalise).
from huggingface_hub import hf_hub_download
from PIL import Image
import importlib.util, torch
spec = importlib.util.spec_from_file_location(
"rvlcdip_models", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_models.py"))
lib = importlib.util.module_from_spec(spec); spec.loader.exec_module(lib)
model, cfg = lib.load("resnet50", "labels-corrected_codes-removed")
x = lib.preprocess(Image.open("page.png"), cfg).unsqueeze(0)
with torch.no_grad():
print(cfg["id2label"][str(model(x).argmax(1).item())])
The DiT-base folders are also in the transformers layout and load with
AutoModelForImageClassification.from_pretrained(..., subfolder="dit_base/labels-corrected_codes-removed").
Each folder's config.json (rvlcdip_config.json for DiT) records the label map, the
normalisation, the selected epoch, and the accuracies above.
Linear SVM baselines
Three linear baselines trained on the same four configurations (one run each; liblinear is deterministic, so there are no seeds). The TF-IDF models read the Amazon Textract text of the page, not the image. Test accuracy (%) on the test split of the model's own training condition:
| model | folder | input | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|---|
| TF-IDF + linear SVM | svm_tfidf |
OCR text | 88.95 | 88.78 | 93.40 | 93.35 | 4.1 MB |
| TF-IDF + handwriting tokens + linear SVM | svm_tfidf_hw |
OCR text | 89.77 | 89.72 | 94.65 | 94.64 | 4.1 MB |
| Linear SVM on CLIP ViT-L/14-336 embeddings | svm_clip_vitl14_336 |
image | 87.47 | 87.44 | 93.77 | 93.59 | 0.1 MB |
- TF-IDF. Word uni- and bigrams, 50,000 features, sublinear term frequency, one-vs-rest
LinearSVC, withCchosen on validation accuracy. - Handwriting tokens. The same model with two pseudo-tokens in front of the text: the number of words on the page in logarithmic bins, and the share of words that Textract marks as handwriting.
- CLIP embeddings. The projected image embedding of the grayscale page from the frozen
openai/clip-vit-large-patch14-336, L2-normalised and standardised, then the same SVM.
The weights are stored as safetensors and the vocabulary as JSON; rvlcdip_svms.py rebuilds the
models with scikit-learn and no pickled objects.
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location(
"rvlcdip_svms", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_svms.py"))
svms = importlib.util.module_from_spec(spec); spec.loader.exec_module(svms)
text_svm = svms.load_tfidf_svm("svm_tfidf", "labels-corrected_codes-removed")
print(text_svm.predict(["Dear Mr. Smith, thank you for your letter of March 3 ..."]))
image_svm = svms.load_clip_svm("labels-corrected_codes-removed") # downloads CLIP
# image_svm.predict([PIL.Image.open("page.png")])
For svm_tfidf_hw, put svms.hw_prefix(n_words, n_handwritten_words) (or
svms.hw_prefix_from_textract(response_text)) in front of each page's text. The decision values
are uncalibrated one-vs-rest margins.
Training
- Data. RVL-CDIP train split (319,999 pages; 296,605 with the corrected labels), model selected by validation accuracy after every epoch.
- CNNs. ImageNet-pretrained torchvision models, SGD (learning rate 0.001, momentum 0.9, polynomial decay), batch 64, 30 epochs, cross-entropy.
- DiT-base.
microsoft/dit-base, AdamW (learning rate 3e-5, cosine schedule with warm-up), batch 32, 10 epochs. - Input. Grayscale page resized to 224 Γ 224. At this size the identification codes are two or three pixels tall.
Each exported file was checked by reloading it through rvlcdip_models.py and comparing its
predictions on 512 test pages with those of the original checkpoint (agreement 99.8β100%).
The SVM exports were checked the same way through rvlcdip_svms.py (identical predictions).
Limitations
- Trained and evaluated on tobacco-litigation scans only; accuracy on other document sources is lower (see the RVL-CDIP-N column).
- The
codes-keptmodels can rely on the identification codes and lose accuracy when the codes are absent. - RVL-CDIP has substantial overlap between its train and test splits, which inflates all test accuracies here.
- GoogLeNet must be built with
transform_input=Trueand without auxiliary heads, as the loader does.
Licence
The weights, configuration files, and loader code in this repository are released under the MIT licence.