RVL-CDIP document classifiers, with and without identification codes

Seven image classifiers and three linear SVM baselines trained on RVL-CDIP (16 document categories), each in four versions that differ in the labels and in the images they were trained on:

folder labels training images
labels-original_codes-kept original RVL-CDIP labels de-identified pages
labels-original_codes-removed original RVL-CDIP labels de-identified pages with the identification codes whitened out
labels-corrected_codes-kept corrected labels de-identified pages
labels-corrected_codes-removed corrected labels de-identified pages with the identification codes whitened out

Identification codes are the Bates numbers stamped on most pages; they are associated with the category label and act as a shortcut (see stefan-hf/rvlcdip-id-codes). The codes were located with stefan-hf/yolov8n-rvlcdip-idcodes. All models were trained on a de-identified copy of the corpus, in which personal data had been replaced with synthetic values, so none of them was trained on the original personal data.

The codes-removed models cannot use the codes, at a cost of 0.0–0.3 points in-domain on the corrected labels. On RVL-CDIP-N, which was collected from other sources, the codes-kept models are 1–6 points more accurate (second table below), so neither version is better on every measure.

Models and accuracy

Test accuracy (%) of the released checkpoint on the test split of its own training condition; in parentheses, the mean over three training seeds. The released checkpoint is always seed 42, the first run, chosen in advance and not by score. Models trained on the original labels are scored on the original test labels (39,999 pages) and models trained on the corrected labels on the corrected test labels (36,446 pages), so the two halves of the table are not directly comparable.

model folder original, codes kept original, codes removed corrected, codes kept corrected, codes removed size
AlexNet alexnet 89.21 (89.30) 88.41 (88.52) 92.85 (92.82) 92.73 (92.73) 228 MB
GoogLeNet googlenet 89.34 (89.25) 88.52 (88.50) 93.19 (93.15) 92.91 (92.97) 23 MB
ResNet-50 resnet50 91.01 (90.85) 90.21 (90.20) 94.26 (94.31) 94.28 (94.19) 94 MB
ResNeXt-50 (32x4d) resnext50 91.48 (91.33) 90.48 (90.63) 94.64 (94.69) 94.56 (94.52) 92 MB
SqueezeNet 1.0 squeezenet 88.70 (88.66) 87.74 (87.69) 92.46 (92.59) 92.50 (92.46) 3 MB
VGG-16 vgg16 91.18 (91.27) 90.74 (90.81) 94.61 (94.62) 94.46 (94.42) 537 MB
DiT-base dit_base 92.69 (92.76) 92.20 (92.19) 95.94 (95.92) 95.74 (95.81) 343 MB

Corrected-label models on both test conditions and on RVL-CDIP-N (1,002 newly collected pages):

model kept β†’ kept kept β†’ removed removed β†’ kept removed β†’ removed RVL-CDIP-N, kept RVL-CDIP-N, removed
AlexNet 92.85 91.54 92.37 92.73 73.6 69.7
GoogLeNet 93.19 92.62 92.72 92.91 77.9 71.5
ResNet-50 94.26 93.71 94.19 94.28 79.2 77.0
ResNeXt-50 (32x4d) 94.64 94.10 94.45 94.56 80.8 76.2
SqueezeNet 1.0 92.46 91.57 92.32 92.50 77.0 75.0
VGG-16 94.61 93.75 94.25 94.46 78.5 77.8
DiT-base 95.94 95.37 95.71 95.74 87.1 83.9

"kept β†’ removed" is a model trained with codes and tested on pages without them.

Usage

rvlcdip_models.py in this repository builds the architecture, loads the weights, and applies the training-time preprocessing (grayscale, bilinear resize to 224 Γ— 224, replicate to three channels, normalise).

from huggingface_hub import hf_hub_download
from PIL import Image
import importlib.util, torch

spec = importlib.util.spec_from_file_location(
    "rvlcdip_models", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_models.py"))
lib = importlib.util.module_from_spec(spec); spec.loader.exec_module(lib)

model, cfg = lib.load("resnet50", "labels-corrected_codes-removed")
x = lib.preprocess(Image.open("page.png"), cfg).unsqueeze(0)
with torch.no_grad():
    print(cfg["id2label"][str(model(x).argmax(1).item())])

The DiT-base folders are also in the transformers layout and load with AutoModelForImageClassification.from_pretrained(..., subfolder="dit_base/labels-corrected_codes-removed"). Each folder's config.json (rvlcdip_config.json for DiT) records the label map, the normalisation, the selected epoch, and the accuracies above.

Linear SVM baselines

Three linear baselines trained on the same four configurations (one run each; liblinear is deterministic, so there are no seeds). The TF-IDF models read the Amazon Textract text of the page, not the image. Test accuracy (%) on the test split of the model's own training condition:

model folder input original, codes kept original, codes removed corrected, codes kept corrected, codes removed size
TF-IDF + linear SVM svm_tfidf OCR text 88.95 88.78 93.40 93.35 4.1 MB
TF-IDF + handwriting tokens + linear SVM svm_tfidf_hw OCR text 89.77 89.72 94.65 94.64 4.1 MB
Linear SVM on CLIP ViT-L/14-336 embeddings svm_clip_vitl14_336 image 87.47 87.44 93.77 93.59 0.1 MB
  • TF-IDF. Word uni- and bigrams, 50,000 features, sublinear term frequency, one-vs-rest LinearSVC, with C chosen on validation accuracy.
  • Handwriting tokens. The same model with two pseudo-tokens in front of the text: the number of words on the page in logarithmic bins, and the share of words that Textract marks as handwriting.
  • CLIP embeddings. The projected image embedding of the grayscale page from the frozen openai/clip-vit-large-patch14-336, L2-normalised and standardised, then the same SVM.

The weights are stored as safetensors and the vocabulary as JSON; rvlcdip_svms.py rebuilds the models with scikit-learn and no pickled objects.

from huggingface_hub import hf_hub_download
import importlib.util

spec = importlib.util.spec_from_file_location(
    "rvlcdip_svms", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_svms.py"))
svms = importlib.util.module_from_spec(spec); spec.loader.exec_module(svms)

text_svm = svms.load_tfidf_svm("svm_tfidf", "labels-corrected_codes-removed")
print(text_svm.predict(["Dear Mr. Smith, thank you for your letter of March 3 ..."]))

image_svm = svms.load_clip_svm("labels-corrected_codes-removed")   # downloads CLIP
# image_svm.predict([PIL.Image.open("page.png")])

For svm_tfidf_hw, put svms.hw_prefix(n_words, n_handwritten_words) (or svms.hw_prefix_from_textract(response_text)) in front of each page's text. The decision values are uncalibrated one-vs-rest margins.

Training

  • Data. RVL-CDIP train split (319,999 pages; 296,605 with the corrected labels), model selected by validation accuracy after every epoch.
  • CNNs. ImageNet-pretrained torchvision models, SGD (learning rate 0.001, momentum 0.9, polynomial decay), batch 64, 30 epochs, cross-entropy.
  • DiT-base. microsoft/dit-base, AdamW (learning rate 3e-5, cosine schedule with warm-up), batch 32, 10 epochs.
  • Input. Grayscale page resized to 224 Γ— 224. At this size the identification codes are two or three pixels tall.

Each exported file was checked by reloading it through rvlcdip_models.py and comparing its predictions on 512 test pages with those of the original checkpoint (agreement 99.8–100%). The SVM exports were checked the same way through rvlcdip_svms.py (identical predictions).

Limitations

  • Trained and evaluated on tobacco-litigation scans only; accuracy on other document sources is lower (see the RVL-CDIP-N column).
  • The codes-kept models can rely on the identification codes and lose accuracy when the codes are absent.
  • RVL-CDIP has substantial overlap between its train and test splits, which inflates all test accuracies here.
  • GoogLeNet must be built with transform_input=True and without auxiliary heads, as the loader does.

Licence

The weights, configuration files, and loader code in this repository are released under the MIT licence.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train stefan-hf/rvlcdip-classifiers