OffLeaf resnet50 (baseline)
Research baseline. Not deployment-ready. See Intended use.
A resnet50 trained on PlantVillage as part of OffLeaf, a study of how far plant-disease classifiers rely on image background rather than on the disease, and whether that reliance can be trained away.
- Repository: https://github.com/AHSharan/OffLeaf
- Training regime:
baseline - Classes: 38
- Seed: 0
The headline
This model scores 99.8% on PlantVillage and 19.5% on real field photographs of the same diseases. That gap, not the lab number, is what the OffLeaf project measures.
Results
Lab accuracy (PlantVillage held-out test)
| metric | value |
|---|---|
| best val accuracy | 0.9980 |
| test accuracy | 0.9978 [0.9965, 0.9991] |
| test macro F1 | 0.9968 |
Field accuracy (zero-shot transfer)
| dataset | accuracy | macro F1 | n |
|---|---|---|---|
| PlantDoc (field) | 0.1953 [0.1802, 0.2101] | 0.1380 | 2580 |
Zero-shot means the model saw no field images during training. This is the setting that matters if a model is ever pointed at a real plant.
Where the model looks
Attribution mass outside the leaf, measured against SAM leaf masks.
gradcam
| metric | value |
|---|---|
offleaf_mass (mass off the leaf) |
0.4920 [0.4767, 0.5047] |
Background-swap benchmark
Same leaf, different background. A model reading the leaf should be unaffected.
| cell | accuracy | n |
|---|---|---|
lab_plain |
0.8650 | 200 |
lab_field |
0.4050 | 200 |
field_plain |
0.1700 | 200 |
field_field |
0.1300 | 200 |
paste_control |
0.9100 | 200 |
Background accounts for 62.6% of the lab-to-field accuracy gap.
Crop and class coverage
- 38 PlantVillage classes across 14 crops: apple, blueberry, cherry, corn, grape, orange, peach, pepper_bell, potato, raspberry, soybean, squash, strawberry, tomato.
- 28 classes have a PlantDoc (field) counterpart - only these can be scored zero-shot.
- 19 classes have PlantSeg lesion-mask coverage. Classes without it cannot be pseudo-labelled, so lesion-level metrics do not apply to them.
Known failure modes
- Collapses on field images. Zero-shot accuracy 19.5% [18.0%, 21.0%] versus near-ceiling in the lab. Treat any field prediction as unreliable.
- Attends to background. 49.2% of Grad-CAM attribution mass falls off the leaf on field images.
- Prediction changes when only the background changes. With the leaf pixel-identical, accuracy moves 86.5% -> 40.5% when the background is swapped from lab to field.
- 63% of the lab-to-field gap is attributable to background rather than to the leaf.
- No abstention. The model always returns a class, with no calibrated way to signal that an image is out of distribution.
- Class coverage is partial. Classes absent from field datasets are untested outside the lab entirely.
Intended use
Intended: research on shortcut learning, background reliance and lab-to-field generalization in plant disease classification; a baseline to compare interventions against; a scoring target for the OffLeaf benchmark.
Not intended: diagnosing disease on real plants, agronomic advice, or any decision affecting a crop. Zero-shot field accuracy is far below what a deployed system would require, and the model has no way to say "I don't know".
A deployable system would need field training data, a leaf detector in front of the classifier, and an abstention option. None of those are here.
How to use
import torch
from models.build import build_model # from github.com/AHSharan/OffLeaf
ck = torch.load("checkpoint.pt", map_location="cpu", weights_only=False)
model = build_model(ck["model_name"], num_classes=ck["num_classes"], pretrained=False)
model.load_state_dict(ck["model_state"])
model.eval()
logits, features = model(images) # features are layer-3 maps, for Grad-CAM
Input: 224x224 RGB, ImageNet normalisation.
Forward returns (logits, feature_map) - the feature map is what the
attribution metrics and the training penalty both differentiate through.
Training
- Backbone:
resnet50, ImageNet-pretrained, single linear head - Optimiser: adamw, lr 0.0003, weight decay 0.0001
- 10 epochs, batch 32, 224px, cosine schedule
- Mixed precision: True
- Training time: 92.8 min
Citation
If you use this model or the OffLeaf benchmark, please also cite the work it builds on: Noyan (2022) for background bias in PlantVillage, and PlantSeg (2024) for the lesion masks the segmenter was trained on.
Card generated 2026-10-03 by release/push_hf.py from the run's own eval
outputs. Every number above traces to a JSON file in the repository.