Instructions to use electblake/clothing_type_classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use electblake/clothing_type_classifier with timm:
import timm model = timm.create_model("hf-hub:electblake/clothing_type_classifier", pretrained=True) - Notebooks
- Google Colab
- Kaggle
Publish three-class clothing type classifier
Browse files- .gitignore +2 -0
- README.md +174 -0
- config.json +50 -0
- preprocessor_config.json +28 -0
.gitignore
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.onnx
|
| 2 |
+
*.safetensors
|
README.md
CHANGED
|
@@ -1,3 +1,177 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
library_name: timm
|
| 4 |
+
pipeline_tag: image-classification
|
| 5 |
+
base_model: timm/tf_efficientnetv2_s.in21k_ft_in1k
|
| 6 |
+
base_model_relation: finetune
|
| 7 |
+
model_name: Clothing Type EfficientNetV2-S
|
| 8 |
+
metrics:
|
| 9 |
+
- accuracy
|
| 10 |
+
tags:
|
| 11 |
+
- image-classification
|
| 12 |
+
- timm
|
| 13 |
+
- pytorch
|
| 14 |
+
- efficientnetv2
|
| 15 |
+
- onnx
|
| 16 |
+
- safetensors
|
| 17 |
+
- clothing-type-classification
|
| 18 |
---
|
| 19 |
+
|
| 20 |
+
# Clothing Type EfficientNetV2-S
|
| 21 |
+
|
| 22 |
+
This is a three-class clothing-type image classifier fine-tuned from
|
| 23 |
+
[`timm/tf_efficientnetv2_s.in21k_ft_in1k`](https://huggingface.co/timm/tf_efficientnetv2_s.in21k_ft_in1k).
|
| 24 |
+
It predicts one of the following labels, in model-output order:
|
| 25 |
+
|
| 26 |
+
| ID | Label |
|
| 27 |
+
|---:|---|
|
| 28 |
+
| 0 | `denim_shorts` |
|
| 29 |
+
| 1 | `lingerie` |
|
| 30 |
+
| 2 | `swim_bikini` |
|
| 31 |
+
|
| 32 |
+
## Model details
|
| 33 |
+
|
| 34 |
+
| Property | Value |
|
| 35 |
+
|---|---|
|
| 36 |
+
| Task | Image classification |
|
| 37 |
+
| Architecture | EfficientNetV2-S |
|
| 38 |
+
| Base model | `timm/tf_efficientnetv2_s.in21k_ft_in1k` |
|
| 39 |
+
| Relationship | Full fine-tune |
|
| 40 |
+
| Precision | FP32; not quantized |
|
| 41 |
+
| Parameters | 20,181,331 |
|
| 42 |
+
| Input | RGB image tensor, `N × 3 × 384 × 256` |
|
| 43 |
+
| Output | Three logits in the label order above |
|
| 44 |
+
| ONNX opset | 18 |
|
| 45 |
+
| License | MIT; the base model is Apache-2.0 |
|
| 46 |
+
|
| 47 |
+
The pretrained 1,000-class head is replaced by a dropout layer with
|
| 48 |
+
probability `0.15` and a three-output linear classifier. The published
|
| 49 |
+
artifacts are full-precision exports, not quantized variants. The ONNX model
|
| 50 |
+
supports dynamic batch sizes and fixes each image to 384 pixels high by 256
|
| 51 |
+
pixels wide.
|
| 52 |
+
|
| 53 |
+
## Files
|
| 54 |
+
|
| 55 |
+
- `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx` is the portable
|
| 56 |
+
inference graph.
|
| 57 |
+
- `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors` contains the
|
| 58 |
+
PyTorch `EfficientNetClassifier` state dictionary.
|
| 59 |
+
- `config.json` records the architecture, labels, and timm model settings.
|
| 60 |
+
- `preprocessor_config.json` records the image preprocessing contract.
|
| 61 |
+
|
| 62 |
+
The safetensors keys include the custom wrapper's `model.` prefix and custom
|
| 63 |
+
classifier head. Load them with the `EfficientNetClassifier` implementation
|
| 64 |
+
from the [source repository](https://github.com/mr-szgz/clothing-classification-model),
|
| 65 |
+
not as an unmodified upstream timm checkpoint.
|
| 66 |
+
|
| 67 |
+
## Preprocessing
|
| 68 |
+
|
| 69 |
+
For the model tensor itself:
|
| 70 |
+
|
| 71 |
+
1. Convert the image to RGB.
|
| 72 |
+
2. Resize it with preserved aspect ratio to fit within `256 × 384`, using
|
| 73 |
+
bicubic interpolation when enlarging and area interpolation when reducing.
|
| 74 |
+
3. Center-pad it to `256 × 384` with RGB value `(127, 127, 127)`.
|
| 75 |
+
4. Rescale unsigned 8-bit pixels by `1 / 255`.
|
| 76 |
+
5. Normalize each channel with mean `0.5` and standard deviation `0.5`.
|
| 77 |
+
6. Convert from HWC to NCHW layout and add the batch dimension.
|
| 78 |
+
|
| 79 |
+
The source data pipeline first detects the largest person with YOLO and crops
|
| 80 |
+
to that bounding box. Inputs that are already tightly person-centered are
|
| 81 |
+
closest to the training distribution. Person detection is not part of the
|
| 82 |
+
published classifier and is unnecessary when the input is already cropped.
|
| 83 |
+
|
| 84 |
+
## ONNX inference
|
| 85 |
+
|
| 86 |
+
Install `huggingface_hub`, `onnxruntime`, `numpy`, `opencv-python-headless`, and
|
| 87 |
+
`Pillow`, then run:
|
| 88 |
+
|
| 89 |
+
```python
|
| 90 |
+
from huggingface_hub import hf_hub_download
|
| 91 |
+
import cv2
|
| 92 |
+
import numpy as np
|
| 93 |
+
import onnxruntime as ort
|
| 94 |
+
from PIL import Image, ImageOps
|
| 95 |
+
|
| 96 |
+
REPO_ID = "electblake/clothing_type_classifier"
|
| 97 |
+
LABELS = ["denim_shorts", "lingerie", "swim_bikini"]
|
| 98 |
+
HEIGHT = 384
|
| 99 |
+
WIDTH = 256
|
| 100 |
+
|
| 101 |
+
model_path = hf_hub_download(
|
| 102 |
+
REPO_ID,
|
| 103 |
+
"clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx",
|
| 104 |
+
)
|
| 105 |
+
|
| 106 |
+
image = ImageOps.exif_transpose(Image.open("person.jpg")).convert("RGB")
|
| 107 |
+
image = np.asarray(image)
|
| 108 |
+
scale = min(HEIGHT / image.shape[0], WIDTH / image.shape[1])
|
| 109 |
+
resized_width = round(image.shape[1] * scale)
|
| 110 |
+
resized_height = round(image.shape[0] * scale)
|
| 111 |
+
interpolation = cv2.INTER_AREA if scale < 1 else cv2.INTER_CUBIC
|
| 112 |
+
image = cv2.resize(image, (resized_width, resized_height), interpolation=interpolation)
|
| 113 |
+
|
| 114 |
+
canvas = np.full((HEIGHT, WIDTH, 3), 127, dtype=np.uint8)
|
| 115 |
+
top = (HEIGHT - resized_height) // 2
|
| 116 |
+
left = (WIDTH - resized_width) // 2
|
| 117 |
+
canvas[top : top + resized_height, left : left + resized_width] = image
|
| 118 |
+
|
| 119 |
+
image = canvas.astype(np.float32) / 255.0
|
| 120 |
+
image = (image - 0.5) / 0.5
|
| 121 |
+
image = np.transpose(image, (2, 0, 1))[None, ...]
|
| 122 |
+
|
| 123 |
+
session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
|
| 124 |
+
logits = session.run(["logits"], {"image": image})[0][0]
|
| 125 |
+
probabilities = np.exp(logits - logits.max())
|
| 126 |
+
probabilities /= probabilities.sum()
|
| 127 |
+
|
| 128 |
+
prediction = int(probabilities.argmax())
|
| 129 |
+
print(LABELS[prediction], float(probabilities[prediction]))
|
| 130 |
+
```
|
| 131 |
+
|
| 132 |
+
## Safetensors inference
|
| 133 |
+
|
| 134 |
+
From a checkout of the source repository:
|
| 135 |
+
|
| 136 |
+
```python
|
| 137 |
+
from huggingface_hub import hf_hub_download
|
| 138 |
+
from safetensors.torch import load_file
|
| 139 |
+
|
| 140 |
+
from src.model import EfficientNetClassifier
|
| 141 |
+
|
| 142 |
+
weights_path = hf_hub_download(
|
| 143 |
+
"electblake/clothing_type_classifier",
|
| 144 |
+
"clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors",
|
| 145 |
+
)
|
| 146 |
+
model = EfficientNetClassifier(
|
| 147 |
+
model_name="tf_efficientnetv2_s.in21k_ft_in1k",
|
| 148 |
+
num_classes=3,
|
| 149 |
+
dropout=0.15,
|
| 150 |
+
pretrained=False,
|
| 151 |
+
)
|
| 152 |
+
model.load_state_dict(load_file(weights_path))
|
| 153 |
+
model.eval()
|
| 154 |
+
```
|
| 155 |
+
|
| 156 |
+
## Training and evaluation
|
| 157 |
+
|
| 158 |
+
The published checkpoint was selected at epoch 29 with **94.84% validation
|
| 159 |
+
accuracy**. This value comes from the checkpoint metadata and is a
|
| 160 |
+
model-selection result on the project's validation split, not an independent
|
| 161 |
+
benchmark or test-set result.
|
| 162 |
+
|
| 163 |
+
Training used weighted cross-entropy, AdamW, MixUp, CutMix, horizontal flips,
|
| 164 |
+
geometric transforms, brightness and contrast changes, and hue and saturation
|
| 165 |
+
changes. The backbone was initially frozen and then fully fine-tuned.
|
| 166 |
+
|
| 167 |
+
## Intended use and limitations
|
| 168 |
+
|
| 169 |
+
The model is intended for research, media organization, and non-critical
|
| 170 |
+
clothing-type tagging. It is not intended for decisions affecting a person's
|
| 171 |
+
rights or access to services.
|
| 172 |
+
|
| 173 |
+
Performance is expected to vary with occlusion, multiple people, unusual
|
| 174 |
+
framing, small subjects, lighting, image quality, and clothing outside the
|
| 175 |
+
three published categories. The classifier always selects one of its three
|
| 176 |
+
known labels and has no `other` or rejection class. Confidence scores are
|
| 177 |
+
softmax probabilities and have not been calibrated.
|
config.json
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architecture": "tf_efficientnetv2_s",
|
| 3 |
+
"num_classes": 3,
|
| 4 |
+
"num_features": 1280,
|
| 5 |
+
"id2label": {
|
| 6 |
+
"0": "denim_shorts",
|
| 7 |
+
"1": "lingerie",
|
| 8 |
+
"2": "swim_bikini"
|
| 9 |
+
},
|
| 10 |
+
"label2id": {
|
| 11 |
+
"denim_shorts": 0,
|
| 12 |
+
"lingerie": 1,
|
| 13 |
+
"swim_bikini": 2
|
| 14 |
+
},
|
| 15 |
+
"pretrained_cfg": {
|
| 16 |
+
"tag": "in21k_ft_in1k",
|
| 17 |
+
"custom_load": false,
|
| 18 |
+
"input_size": [
|
| 19 |
+
3,
|
| 20 |
+
384,
|
| 21 |
+
256
|
| 22 |
+
],
|
| 23 |
+
"fixed_input_size": true,
|
| 24 |
+
"interpolation": "bicubic",
|
| 25 |
+
"resize_mode": "longest_max_size",
|
| 26 |
+
"padding": "center",
|
| 27 |
+
"pad_value": [
|
| 28 |
+
127,
|
| 29 |
+
127,
|
| 30 |
+
127
|
| 31 |
+
],
|
| 32 |
+
"mean": [
|
| 33 |
+
0.5,
|
| 34 |
+
0.5,
|
| 35 |
+
0.5
|
| 36 |
+
],
|
| 37 |
+
"std": [
|
| 38 |
+
0.5,
|
| 39 |
+
0.5,
|
| 40 |
+
0.5
|
| 41 |
+
],
|
| 42 |
+
"num_classes": 3,
|
| 43 |
+
"pool_size": [
|
| 44 |
+
12,
|
| 45 |
+
8
|
| 46 |
+
],
|
| 47 |
+
"first_conv": "conv_stem",
|
| 48 |
+
"classifier": "classifier"
|
| 49 |
+
}
|
| 50 |
+
}
|
preprocessor_config.json
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architecture": "tf_efficientnetv2_s",
|
| 3 |
+
"data_config": {
|
| 4 |
+
"input_size": [
|
| 5 |
+
3,
|
| 6 |
+
384,
|
| 7 |
+
256
|
| 8 |
+
],
|
| 9 |
+
"interpolation": "bicubic",
|
| 10 |
+
"mean": [
|
| 11 |
+
0.5,
|
| 12 |
+
0.5,
|
| 13 |
+
0.5
|
| 14 |
+
],
|
| 15 |
+
"std": [
|
| 16 |
+
0.5,
|
| 17 |
+
0.5,
|
| 18 |
+
0.5
|
| 19 |
+
],
|
| 20 |
+
"resize_mode": "longest_max_size",
|
| 21 |
+
"padding": "center",
|
| 22 |
+
"pad_value": [
|
| 23 |
+
127,
|
| 24 |
+
127,
|
| 25 |
+
127
|
| 26 |
+
]
|
| 27 |
+
}
|
| 28 |
+
}
|