DINO Plankton Classifier
The model is trained on PML IFCB data consisting of 145 plankton classes. The DINOv3 backbone is frozen and only the MLP head is trained, for 20 epochs, with classes under 100 training images oversampled to 100 by duplication. On the 14,837-image PML IFCB test split it reaches 0.941 accuracy and 0.874 macro-F1.
Inference
Use the provided inference script. See example in demo.ipynb on predicting the classes for two synthetic samples.
CLIP-LoRA Classifier
The repo also includes a second classifier for the same 145 classes, built on OpenAI CLIP ViT-B/16. It was fine-tuned with rank-32 LoRA and a hierarchical taxonomy-contrastive loss on Planktonzilla. This is the encoder whose embeddings condition the CLIP variant of FineDiffusion in our ECCV MARINE Workshop paper. Its files are clip_model.py, clip_inference.py, clip_config.json and clip_weights.pt. clip_weights.pt (16 MB) holds only the LoRA adapters, the retrained final LayerNorms and the class prototypes. The backbone downloads through open_clip, and the adapters are merged into it at load time.
The classifier compares an image's embedding with one prototype per class and returns the most similar class. model.predict(pixel_values, prototypes=...) chooses which prototypes to use:
"image"(default) uses class-mean image prototypes. Each class prototype is the average embedding of that class's labelled IFCB training images. No training is involved, only averaging. The question it answers is "which class's images does this image look most like?""text"gives zero-shot classification. Each class prototype is the text embedding of the class's taxonomic lineage, e.g. "Chromista Heterokontophyta Bacillariophyceae ... Ditylum brightwellii". It needs no labelled images and works for any string, but it depends on whether the classes were in the Planktonzilla training set. There is a tradeoff between generalisation across instruments and species, and performance on the PML IFCB dataset specifically.
prototypes |
IFCB test accuracy | macro-F1 |
|---|---|---|
"image" |
0.773 | 0.636 |
"text" |
0.282 | 0.145 |
These results come from the 14,837-image PML IFCB test split. The CLIP model never saw PML data during training (Planktonzilla includes IFCB images from WHOI and SYKE, but not PML), and the DINO classifier above reaches 0.874 macro-F1 on the same split. The softmax probabilities are not calibrated. See demo_clip.ipynb.
- Downloads last month
- 14
Model tree for danielaivanova/dino_plankton_classifier
Base model
facebook/dinov3-vit7b16-pretrain-lvd1689m