OffLeaf resnet50 (baseline)

Research baseline. Not deployment-ready. See Intended use.

A resnet50 trained on PlantVillage as part of OffLeaf, a study of how far plant-disease classifiers rely on image background rather than on the disease, and whether that reliance can be trained away.

The headline

This model scores 99.8% on PlantVillage and 19.5% on real field photographs of the same diseases. That gap, not the lab number, is what the OffLeaf project measures.

Results

Lab accuracy (PlantVillage held-out test)

metric value
best val accuracy 0.9980
test accuracy 0.9978 [0.9965, 0.9991]
test macro F1 0.9968

Field accuracy (zero-shot transfer)

dataset accuracy macro F1 n
PlantDoc (field) 0.1953 [0.1802, 0.2101] 0.1380 2580

Zero-shot means the model saw no field images during training. This is the setting that matters if a model is ever pointed at a real plant.

Where the model looks

Attribution mass outside the leaf, measured against SAM leaf masks.

gradcam

metric value
offleaf_mass (mass off the leaf) 0.4920 [0.4767, 0.5047]

Background-swap benchmark

Same leaf, different background. A model reading the leaf should be unaffected.

cell accuracy n
lab_plain 0.8650 200
lab_field 0.4050 200
field_plain 0.1700 200
field_field 0.1300 200
paste_control 0.9100 200

Background accounts for 62.6% of the lab-to-field accuracy gap.

Crop and class coverage

  • 38 PlantVillage classes across 14 crops: apple, blueberry, cherry, corn, grape, orange, peach, pepper_bell, potato, raspberry, soybean, squash, strawberry, tomato.
  • 28 classes have a PlantDoc (field) counterpart - only these can be scored zero-shot.
  • 19 classes have PlantSeg lesion-mask coverage. Classes without it cannot be pseudo-labelled, so lesion-level metrics do not apply to them.

Known failure modes

  • Collapses on field images. Zero-shot accuracy 19.5% [18.0%, 21.0%] versus near-ceiling in the lab. Treat any field prediction as unreliable.
  • Attends to background. 49.2% of Grad-CAM attribution mass falls off the leaf on field images.
  • Prediction changes when only the background changes. With the leaf pixel-identical, accuracy moves 86.5% -> 40.5% when the background is swapped from lab to field.
  • 63% of the lab-to-field gap is attributable to background rather than to the leaf.
  • No abstention. The model always returns a class, with no calibrated way to signal that an image is out of distribution.
  • Class coverage is partial. Classes absent from field datasets are untested outside the lab entirely.

Intended use

Intended: research on shortcut learning, background reliance and lab-to-field generalization in plant disease classification; a baseline to compare interventions against; a scoring target for the OffLeaf benchmark.

Not intended: diagnosing disease on real plants, agronomic advice, or any decision affecting a crop. Zero-shot field accuracy is far below what a deployed system would require, and the model has no way to say "I don't know".

A deployable system would need field training data, a leaf detector in front of the classifier, and an abstention option. None of those are here.

How to use

import torch
from models.build import build_model      # from github.com/AHSharan/OffLeaf

ck = torch.load("checkpoint.pt", map_location="cpu", weights_only=False)
model = build_model(ck["model_name"], num_classes=ck["num_classes"], pretrained=False)
model.load_state_dict(ck["model_state"])
model.eval()

logits, features = model(images)   # features are layer-3 maps, for Grad-CAM

Input: 224x224 RGB, ImageNet normalisation. Forward returns (logits, feature_map) - the feature map is what the attribution metrics and the training penalty both differentiate through.

Training

  • Backbone: resnet50, ImageNet-pretrained, single linear head
  • Optimiser: adamw, lr 0.0003, weight decay 0.0001
  • 10 epochs, batch 32, 224px, cosine schedule
  • Mixed precision: True
  • Training time: 92.8 min

Citation

If you use this model or the OffLeaf benchmark, please also cite the work it builds on: Noyan (2022) for background bias in PlantVillage, and PlantSeg (2024) for the lesion masks the segmenter was trained on.


Card generated 2026-10-03 by release/push_hf.py from the run's own eval outputs. Every number above traces to a JSON file in the repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support