Install the HFTrainer repository before running the commands below. This repository hosts the processed pretrained base; the public-data training outputs are separate. Artifact provenance.

Vision Transformer · ViT-B/16

Image classification with repository-owned model, trainer and inference code.

Verified: 20 real-data training steps, saved checkpoint and native checkpoint-only base inference. Convergence is not established.

All models · Settings · Train · Infer · Evidence · Demos

At a glance

Property Released setting
Model Base · patch 16 · 224 px
Training Full encoder + new three-class head
Training input 224 × 224
Public dataset beans
Runtime Local HFTrainer implementation; supporting PyTorch/media libraries and model assets remain dependencies

Sources

Original paper / report · Original code

The original repository is provenance, not a runtime checkout requirement. Third-party notices preserve implementation and asset terms.

Settings and checkpoints

Setting Training config Processed checkpoint Access
Base · patch 16 · 224 px config.py HFTrainer-ViT-Base-Patch16-224 Public

Only the setting above is a released HFTrainer artifact. Its root config.json, weights and all task-required component configs, processors and schedulers are included together. Inference needs only the checkpoint path or Hub ID, plus task inputs.

Pinned weight revision: 18ca4731e4d1. Original conversion evidence describes the initial export; the current revision adds checkpoint-only dispatch metadata without changing tensors.

Setup

Run from the repository root:

python -m pip install -e "."
python -m pip install "huggingface_hub>=0.34,<2"

Verified on one NVIDIA H200 with PyTorch 2.8.0+cu128. This is the tested environment, not a measured minimum-memory requirement.

Data

The recipe downloads ZeyuLing/hftrainer_beans to data/hftrainer_beans. Split counts: 1034 train / 133 validation / 128 test. Dataset source, license, transformations and checksum verification are documented in Public demo datasets.

Model features are cached locally when required. The script never substitutes synthetic samples for missing real media.

Train

One command downloads pinned data and weights, prepares required caches, and runs the 20-step recipe:

python tools/run_public_demo.py vit

Reuse downloaded weights with --checkpoint path/to/complete_bundle; change the output directory with --work-dir path/to/run. The underlying command is python tools/train.py configs/public_data/vit.py. Seed: 42. Training logs and checkpoints are saved under work_dirs/public_data/vit/. If you reused a custom checkpoint directory, set HFTRAINER_CHECKPOINT to that directory before directly invoking training, resume or export commands.

The final resumable checkpoint is work_dirs/public_data/vit/checkpoint-iter_20. For a longer resumed run, keep the same checkpoint/data setting:

python tools/train.py configs/public_data/vit.py --auto-resume \
  --cfg-options train_cfg.max_iters=40

Infer

Run the processed pretrained base, independently of any training config:

python tools/infer.py --model ZeyuLing/HFTrainer-ViT-Base-Patch16-224 \
  --revision 18ca4731e4d1d3f168638a13cd925f759c8cee8c --device cuda --seed 42 \
  --input path/to/image.jpg

--model also accepts a local complete checkpoint directory. A resumable training checkpoint is not a standalone model: export with tools/export_model.py using the matching training config, then pass the exported directory to --model. Artifact and configuration contract.

To package this training run as a complete checkpoint (including its frozen base components):

python tools/export_model.py --config configs/public_data/vit.py \
  --checkpoint work_dirs/public_data/vit/checkpoint-iter_20 --output exports/vit

Then run the inference command above with --model exports/vit and omit --revision. Complete exports duplicate the base weights; allow sufficient disk space.

Evidence and loss

Full pretrained export/reload produced identical logits; tiny numerical reference checks also pass.

Vision Transformer · ViT-B/16 raw 20-step training losses on beans

Raw training record · Training log

These are raw, unsmoothed objectives on the public dataset. Twenty steps establish pipeline execution, not convergence. Loss values are not comparable across models.

Held-out accuracy: 72.18% on all 133 Beans validation images after 20 steps. The 128-image test split was not used to tune this run.

Demos

Held-out Beans classification examples from the trained three-class head

Eight fixed validation examples, labels and predictions from the trained three-class head. The published base uses its original ImageNet head.

Limitations

Only B/16 at 224 px is packaged here. Upstream Large/Huge variants are not released HFTrainer settings. The published base has its original ImageNet head; the Beans recipe replaces it with three classes.

Citation

@misc{dosovitskiy2021vit,
  title = {An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale},
  author = {Alexey Dosovitskiy and others},
  year = {2021},
  eprint = {2010.11929},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2010.11929}
}

Also retain the dataset citation and attribution.

Downloads last month
38
Safetensors
Model size
86.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for ZeyuLing/HFTrainer-ViT-Base-Patch16-224