Image Classification
timm
LiteRT
Safetensors
coffee
plant-disease
agriculture
efficientnet
supervised-fine-tuning

Coffee Leaf EfficientNet-B2

Project repository: Imhaohao/global-ai-hackathon.

B2 is the featured model and the default inference choice. This repository packages the selected eight-class EfficientNet-B2 coffee-leaf classifier, with B0, B1 and B3 retained as comparison variants. B2 has the best aggregate accuracy and macro F1 on the shared internal test set. External rust and scan results remain weak; the model is a research prototype whose internal scores do not establish reliable field diagnosis.

Download B2: PyTorch weights · Mobile TFLite export · Configuration · Acceptance rules.

The original coffee-specific model used for B0 was Huyt's Arabica coffee EfficientNet-B0. B2 itself starts from timm/efficientnet_b2.ra_in1k. B1-B3 learn independently from their ImageNet initializations; they are not scaled copies of Huyt's five-class coffee head. The selected weights in this package are modified fine-tuned derivatives, with full credits in ATTRIBUTION.md.

Internal comparison

All four exported models were compared on the same 5,868 held-out internal images, with eight output classes, 224-pixel inputs, the same brightness/blur policy and separate validation-calibrated confidence thresholds. Metrics below describe the TFLite exports rather than only their research checkpoints.

Model Raw accuracy Macro F1 Accepted accuracy Accepted coverage Mean response ms Export MB
B2 94.09% 0.8720 95.14% 79.52% 98.1 8.57
B0 86.79% 0.7614 88.09% 80.11% 62.7 8.09
B1 93.27% 0.8640 94.16% 79.38% 91.3 7.32
B3 93.22% 0.8544 94.14% 79.11% 125.2 11.80

Accepted accuracy is the fraction correct among accepted diagnoses. Accepted coverage is the fraction of all images accepted, including wrong diagnoses. Raw accuracy scores the highest-probability label before rejection. MB uses 1,000,000 bytes. Mean response is from 96 timed single-image trials per model on a desktop CPU with four threads.

Four-model comparison

B2 has the highest internal aggregate score, but B1 has stronger Cercospora F1. Only 26 internal test examples are labelled red spider mite, so that class's score is especially uncertain. The synthetic smaller-leaf probe is a transformation of held-out crops; it is not an external field study.

Class B2 F1 B0 F1 B1 F1 B3 F1 Test support
Cercospora 0.717 0.677 0.787 0.754 231
Healthy 0.947 0.733 0.955 0.959 318
Leaf miner 0.955 0.900 0.935 0.939 1253
Phoma 0.891 0.818 0.832 0.841 677
Rust 0.954 0.898 0.957 0.956 2267
Red spider mite 0.540 0.319 0.483 0.435 26
Weevil damage 0.976 0.924 0.968 0.961 894
Unsupported 0.998 0.822 0.995 0.990 202

Internal confusion matrices

The comparison measures the delivered systems, not architecture alone: B0 uses older fine-tuning data and coffee-specific initialization, while B1-B3 share expanded data but have different ImageNet pretraining recipes. B0's upstream checkpoint already learned from AGML; the full report includes subsets excluding AGML and known pretraining exposure. Crop/group disjointness cannot guarantee independent plants where plant identifiers are absent.

External rust results

The AgML coffee-rust RGB collection supplied 1,119 evaluated images: 847 Rust and 272 NoRust. None trained or selected these models. One of the original 1,120 images was conservatively excluded as a near-match candidate. NoRust is a rust-negative label and does not establish a healthy leaf or another disease.

Model Correct rust accepted Wrong diagnosis accepted on rust Rust rejected False rust accepted on NoRust AUROC
B2 16 / 847 (1.9%) 177 654 0 / 272 0.778
B0 143 / 847 (16.9%) 30 674 74 / 272 0.492
B1 22 / 847 (2.6%) 169 656 0 / 272 0.742
B3 52 / 847 (6.1%) 114 681 0 / 272 0.697

B2 correctly identified and accepted just 16 of 847 rust images (1.9%). It accepted an incorrect supported diagnosis on 177 rust images and rejected 654. Its highest-scoring label was rust on 44 images; 28 of those were rejected. AUROC evaluates score ranking rather than the final diagnosis decision. A higher AUROC does not mean more correct accepted diagnoses.

External rust decision outcomes

Six B2 accepted rust examples

The six displayed examples were randomly selected from B2's 16 correct accepted successes and reverified with the current export and acceptance rules. This is a success-only sample, not a representative performance sample. Download the exact external test images and B2 results.

External scan results

The separate PG26038 sample bank has 100 scans of 50 paired leaves: 40 rust, 40 Mycena and 20 healthy images. Mycena maps to unsupported. Opposite sides of the same leaf are repeated measures.

Model Raw accuracy Accepted / 100 Correct accepted / 100 Accuracy among accepted
B2 12.0% 89 10 11.2%
B0 27.0% 55 9 16.4%
B1 44.0% 55 19 34.5%
B3 42.0% 63 28 44.4%

The external collections have been inspected during development and should be described as development stress tests rather than newly blinded field tests. None of the four models is established as dependable for these external imaging conditions.

Training data

B2 uses 20,311 prepared training images/crops, with 5,294 validation images. B1 and B3 use exactly the same expanded manifest. Selected B0 uses 14,119 training and 3,921 validation entries. The comparison retains the expanded 5,868-image test set for all four models; B0's original test split has 4,571 entries.

Training source B0 B1 / B2 / B3 each Recorded license
AgML Arabica / JMuBEN 1278 1278 CC BY 4.0
BRACOL whole leaves 964 964 CC BY 4.0 data; MIT code
BRACOL symptom crops 1462 1462 CC BY 4.0 data; MIT code
RoCoLe 1072 1072 CC BY 4.0
CoffeeLeaf-CO field collection 5448 5448 CC BY 4.0
Silva subset through CoffeeLeaf-CO 2861 2861 CC BY 4.0
Makerere Beans 1034 1034 MIT per datasetcard
Saposoa, Peru 0 1105 CC BY 4.0
Expert-reviewed BRACOL crops 0 5087 CC BY 4.0
Total 14,119 20,311 Per source

Training source composition

Counts are prepared image/crop entries, not independent plants or original photographs. BRACOL symptom and reviewed crops can share photographs; their parent/group splits are inherited. The Silva subset is already inside CoffeeLeaf-CO and is not counted as a separately downloaded independent collection. Empty CoffeeLeaf annotations were not relabelled healthy. Beans is explicitly unsupported; Mycena in Peru and the scan bank is also unsupported, not Cercospora.

Exact source, class and split counts, hashes, parent/group identities, licenses and source decisions are available in:

Peru and reviewed BRACOL were admitted for B1-B3. Separate expanded-data B0 experiments failed validation safeguards and did not replace the selected B0. The augmented Uganda source was excluded after unresolved transformed overlap; Xinzhai was downloaded but not used pending label-taxonomy and overlap checks. Other excluded candidates include synthetic coffee images and datasets with missing or incompatible labels. The registry is a list of candidates, not 29 training datasets. Full training-image archives remain with their original publishers; this package contains manifests and source links.

Training procedure

Training used supervised fine-tuning, not reinforcement learning. Source-provided labels supervise class-weighted cross-entropy with 0.05 label smoothing and AdamW. No model-generated labels replace source labels. The classifier head learns first, then selected later backbone layers are fine-tuned at a low learning rate. Validation selects checkpoints and acceptance thresholds; test/external images do not select them.

B0 trained its classifier for 30 epochs and ran three partial-backbone epochs, selecting backbone epoch two. Its classifier learning rate was 0.001 and partial-backbone learning rate 0.00003, with weight decay 0.01. A later distilled/source-balanced B0 candidate was rejected; its teacher regularizer was not reinforcement learning and is not the selected model.

B1-B3 share the regularized recipe: classifier learning rate 0.001, weight decay 0.01, dropout 0.15, maximum 30 epochs, patience five after epoch ten; then last two blocks/head fine-tuning at learning rate 0.00003, weight decay 0.03, frozen BatchNorm, gradient clipping one, maximum six epochs and patience two. Grouped parent rotation and weak-class replay reduce domination by correlated crops. The partial-backbone stage uses 70% canonical views and 30% full-target scale views at 65-100%, with mild color jitter and flips. Context views preserve supplied target regions; the canonical center crop can still trim elongated leaves.

Checkpoint selection combines source-balanced F1 with smaller-leaf and context validation scores, with guard checks on weak classes. The selected B1, B2 and B3 backbone checkpoint is epoch six. Regularization and validation checks limit overfitting risk but do not guarantee external generalization. Detailed selected histories and verification records are under evidence/.

Original model credits

Variant Pinned original checkpoint Upstream training Original model license
B2 timm/efficientnet_b2.ra_in1k ImageNet-1k Apache-2.0
B0 Huyt/arabica-coffee-leaf-disease-efficientnet-b0 ImageNet + prior AGML coffee training CC BY 4.0
B1 timm/efficientnet_b1.ft_in1k ImageNet-1k Apache-2.0
B3 timm/efficientnet_b3.ra2_in1k ImageNet-1k Apache-2.0

Credit the EfficientNet architecture authors Mingxing Tan and Quoc V. Le, timm / Ross Wightman, Huyt, and the dataset authors. See ATTRIBUTION.md for names, DOI links, change notices and BibTeX. Component-specific licenses and complete upstream license texts are in LICENSES.md and licenses/.

Use the models

B2's model.safetensors, corrected eight-class config.json, exact model.tflite export, model-config.json acceptance rules and pinned provenance.json are at the repository root, so B2 is the principal Hugging Face model. B0, B1 and B3 have the same file set in their respective models/ subfolders. B3's 11.80 MB export exceeds the unchanged 10 MB app limit and is a comparison candidate rather than an installed app option.

Install the inference dependencies in a Python environment, then run:

python -m pip install -r requirements.txt
python predict.py leaf.jpg --model b2
python predict.py leaf.jpg --model b2 --backend pytorch

Use --model b0, b1 or b3 for a comparison variant. The default TFLite backend takes NHWC float32 RGB values in 0-255 and performs normalization and temperature calibration inside the export. It applies the documented desktop blur, brightness and per-class acceptance policy. The PyTorch backend normalizes using the packaged mean/std and applies temperature outside the checkpoint; export compression can produce small differences from the float32 checkpoint.

Canonical preprocessing resizes the short side to 256 with Pillow bicubic and center-crops to 224; it uses crop_pct=0.875, mean [0.485, 0.456, 0.406], and std [0.229, 0.224, 0.225]. This is the fine-tuned project's contract for every variant, not the original timm B2/B3 published test resolution. Brightness limits come from the original supported training set and remain shared. An accepted label or high model score does not establish diagnostic correctness; unsupported and disabled classes are rejected.

The root B2 checkpoint follows timm's Hugging Face configuration format; comparison variants are in subfolders. This is not a Transformers model or a promise of automatic hosted inference. Run the script locally or integrate a chosen variant using its supplied configuration. No credentials or network access are required for packaged-model inference.

Desktop timing

The four-model benchmark used eight class-balanced images with 12 repetitions each, 96 timed predictions per model, four CPU threads and balanced invocation order. Mean total responses were B0 62.7 ms, B1 91.3 ms, B2 98.1 ms, and B3 125.2 ms. Mean inference-only times were 57.6, 86.1, 92.9 and 120.1 ms, respectively. Models were already loaded, and the operating-system file cache was warm.

A separate fresh B2 pass over all 1,119 external rust test images took 105.15 seconds, averaging 93.9 ms/image and 10.64 images/second. Model inference alone took 99.48 seconds. This one-pass dataset result includes model initialization, local decoding, preprocessing and acceptance decisions; it excludes Python startup, file hashing, report verification and serialization. It is a different input distribution from the repeated eight-image benchmark.

Single-image timing report · Whole-dataset B2 timing · Per-image dataset timings. Desktop CPU results do not measure phone capture, native mobile bridging, network transfer or UI rendering.

Limitations and verification

Internal crop-based scores do not establish full-plant or real-world farm performance. Camera distance, lighting, background, geography, species and capture method differ across sources; external rust/scans show substantial failure. Mites have limited training/test support. Mixed conditions, unseen pests and unsupported crops are outside the supported label taxonomy. Healthy predictions do not establish the absence of every disease.

Models and dataset inputs are frozen; packaging performs no new training or checkpoint selection. File checksums are in SHA256SUMS.txt, and load/inference checks are recorded in verification.json. Evaluation numbers originate in the preserved local training and comparison reports, not copied upstream headline scores. Full comparison data.

Hugging Face upload

This folder is an upload-ready local model-page package. Create a model repository in your Hugging Face account and upload this entire folder, preserving the relative paths. Its root README.md is the model card and credits B2 as the principal base model; each comparison variant also has its own card and original-checkpoint attribution. index.html is an offline preview for the Desktop. The page and files have not been published to a remote account by creating this package.

Downloads last month
29
Safetensors
Model size
7.78M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ConnorLee08/coffee-leaf-efficientnet-b2

Finetuned
(2)
this model

Datasets used to train ConnorLee08/coffee-leaf-efficientnet-b2