electblake commited on
Commit
6fd27bf
·
verified ·
1 Parent(s): cd1b764

Publish three-class clothing type classifier

Browse files
Files changed (4) hide show
  1. .gitignore +2 -0
  2. README.md +174 -0
  3. config.json +50 -0
  4. preprocessor_config.json +28 -0
.gitignore ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ *.onnx
2
+ *.safetensors
README.md CHANGED
@@ -1,3 +1,177 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ library_name: timm
4
+ pipeline_tag: image-classification
5
+ base_model: timm/tf_efficientnetv2_s.in21k_ft_in1k
6
+ base_model_relation: finetune
7
+ model_name: Clothing Type EfficientNetV2-S
8
+ metrics:
9
+ - accuracy
10
+ tags:
11
+ - image-classification
12
+ - timm
13
+ - pytorch
14
+ - efficientnetv2
15
+ - onnx
16
+ - safetensors
17
+ - clothing-type-classification
18
  ---
19
+
20
+ # Clothing Type EfficientNetV2-S
21
+
22
+ This is a three-class clothing-type image classifier fine-tuned from
23
+ [`timm/tf_efficientnetv2_s.in21k_ft_in1k`](https://huggingface.co/timm/tf_efficientnetv2_s.in21k_ft_in1k).
24
+ It predicts one of the following labels, in model-output order:
25
+
26
+ | ID | Label |
27
+ |---:|---|
28
+ | 0 | `denim_shorts` |
29
+ | 1 | `lingerie` |
30
+ | 2 | `swim_bikini` |
31
+
32
+ ## Model details
33
+
34
+ | Property | Value |
35
+ |---|---|
36
+ | Task | Image classification |
37
+ | Architecture | EfficientNetV2-S |
38
+ | Base model | `timm/tf_efficientnetv2_s.in21k_ft_in1k` |
39
+ | Relationship | Full fine-tune |
40
+ | Precision | FP32; not quantized |
41
+ | Parameters | 20,181,331 |
42
+ | Input | RGB image tensor, `N × 3 × 384 × 256` |
43
+ | Output | Three logits in the label order above |
44
+ | ONNX opset | 18 |
45
+ | License | MIT; the base model is Apache-2.0 |
46
+
47
+ The pretrained 1,000-class head is replaced by a dropout layer with
48
+ probability `0.15` and a three-output linear classifier. The published
49
+ artifacts are full-precision exports, not quantized variants. The ONNX model
50
+ supports dynamic batch sizes and fixes each image to 384 pixels high by 256
51
+ pixels wide.
52
+
53
+ ## Files
54
+
55
+ - `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx` is the portable
56
+ inference graph.
57
+ - `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors` contains the
58
+ PyTorch `EfficientNetClassifier` state dictionary.
59
+ - `config.json` records the architecture, labels, and timm model settings.
60
+ - `preprocessor_config.json` records the image preprocessing contract.
61
+
62
+ The safetensors keys include the custom wrapper's `model.` prefix and custom
63
+ classifier head. Load them with the `EfficientNetClassifier` implementation
64
+ from the [source repository](https://github.com/mr-szgz/clothing-classification-model),
65
+ not as an unmodified upstream timm checkpoint.
66
+
67
+ ## Preprocessing
68
+
69
+ For the model tensor itself:
70
+
71
+ 1. Convert the image to RGB.
72
+ 2. Resize it with preserved aspect ratio to fit within `256 × 384`, using
73
+ bicubic interpolation when enlarging and area interpolation when reducing.
74
+ 3. Center-pad it to `256 × 384` with RGB value `(127, 127, 127)`.
75
+ 4. Rescale unsigned 8-bit pixels by `1 / 255`.
76
+ 5. Normalize each channel with mean `0.5` and standard deviation `0.5`.
77
+ 6. Convert from HWC to NCHW layout and add the batch dimension.
78
+
79
+ The source data pipeline first detects the largest person with YOLO and crops
80
+ to that bounding box. Inputs that are already tightly person-centered are
81
+ closest to the training distribution. Person detection is not part of the
82
+ published classifier and is unnecessary when the input is already cropped.
83
+
84
+ ## ONNX inference
85
+
86
+ Install `huggingface_hub`, `onnxruntime`, `numpy`, `opencv-python-headless`, and
87
+ `Pillow`, then run:
88
+
89
+ ```python
90
+ from huggingface_hub import hf_hub_download
91
+ import cv2
92
+ import numpy as np
93
+ import onnxruntime as ort
94
+ from PIL import Image, ImageOps
95
+
96
+ REPO_ID = "electblake/clothing_type_classifier"
97
+ LABELS = ["denim_shorts", "lingerie", "swim_bikini"]
98
+ HEIGHT = 384
99
+ WIDTH = 256
100
+
101
+ model_path = hf_hub_download(
102
+ REPO_ID,
103
+ "clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx",
104
+ )
105
+
106
+ image = ImageOps.exif_transpose(Image.open("person.jpg")).convert("RGB")
107
+ image = np.asarray(image)
108
+ scale = min(HEIGHT / image.shape[0], WIDTH / image.shape[1])
109
+ resized_width = round(image.shape[1] * scale)
110
+ resized_height = round(image.shape[0] * scale)
111
+ interpolation = cv2.INTER_AREA if scale < 1 else cv2.INTER_CUBIC
112
+ image = cv2.resize(image, (resized_width, resized_height), interpolation=interpolation)
113
+
114
+ canvas = np.full((HEIGHT, WIDTH, 3), 127, dtype=np.uint8)
115
+ top = (HEIGHT - resized_height) // 2
116
+ left = (WIDTH - resized_width) // 2
117
+ canvas[top : top + resized_height, left : left + resized_width] = image
118
+
119
+ image = canvas.astype(np.float32) / 255.0
120
+ image = (image - 0.5) / 0.5
121
+ image = np.transpose(image, (2, 0, 1))[None, ...]
122
+
123
+ session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
124
+ logits = session.run(["logits"], {"image": image})[0][0]
125
+ probabilities = np.exp(logits - logits.max())
126
+ probabilities /= probabilities.sum()
127
+
128
+ prediction = int(probabilities.argmax())
129
+ print(LABELS[prediction], float(probabilities[prediction]))
130
+ ```
131
+
132
+ ## Safetensors inference
133
+
134
+ From a checkout of the source repository:
135
+
136
+ ```python
137
+ from huggingface_hub import hf_hub_download
138
+ from safetensors.torch import load_file
139
+
140
+ from src.model import EfficientNetClassifier
141
+
142
+ weights_path = hf_hub_download(
143
+ "electblake/clothing_type_classifier",
144
+ "clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors",
145
+ )
146
+ model = EfficientNetClassifier(
147
+ model_name="tf_efficientnetv2_s.in21k_ft_in1k",
148
+ num_classes=3,
149
+ dropout=0.15,
150
+ pretrained=False,
151
+ )
152
+ model.load_state_dict(load_file(weights_path))
153
+ model.eval()
154
+ ```
155
+
156
+ ## Training and evaluation
157
+
158
+ The published checkpoint was selected at epoch 29 with **94.84% validation
159
+ accuracy**. This value comes from the checkpoint metadata and is a
160
+ model-selection result on the project's validation split, not an independent
161
+ benchmark or test-set result.
162
+
163
+ Training used weighted cross-entropy, AdamW, MixUp, CutMix, horizontal flips,
164
+ geometric transforms, brightness and contrast changes, and hue and saturation
165
+ changes. The backbone was initially frozen and then fully fine-tuned.
166
+
167
+ ## Intended use and limitations
168
+
169
+ The model is intended for research, media organization, and non-critical
170
+ clothing-type tagging. It is not intended for decisions affecting a person's
171
+ rights or access to services.
172
+
173
+ Performance is expected to vary with occlusion, multiple people, unusual
174
+ framing, small subjects, lighting, image quality, and clothing outside the
175
+ three published categories. The classifier always selects one of its three
176
+ known labels and has no `other` or rejection class. Confidence scores are
177
+ softmax probabilities and have not been calibrated.
config.json ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architecture": "tf_efficientnetv2_s",
3
+ "num_classes": 3,
4
+ "num_features": 1280,
5
+ "id2label": {
6
+ "0": "denim_shorts",
7
+ "1": "lingerie",
8
+ "2": "swim_bikini"
9
+ },
10
+ "label2id": {
11
+ "denim_shorts": 0,
12
+ "lingerie": 1,
13
+ "swim_bikini": 2
14
+ },
15
+ "pretrained_cfg": {
16
+ "tag": "in21k_ft_in1k",
17
+ "custom_load": false,
18
+ "input_size": [
19
+ 3,
20
+ 384,
21
+ 256
22
+ ],
23
+ "fixed_input_size": true,
24
+ "interpolation": "bicubic",
25
+ "resize_mode": "longest_max_size",
26
+ "padding": "center",
27
+ "pad_value": [
28
+ 127,
29
+ 127,
30
+ 127
31
+ ],
32
+ "mean": [
33
+ 0.5,
34
+ 0.5,
35
+ 0.5
36
+ ],
37
+ "std": [
38
+ 0.5,
39
+ 0.5,
40
+ 0.5
41
+ ],
42
+ "num_classes": 3,
43
+ "pool_size": [
44
+ 12,
45
+ 8
46
+ ],
47
+ "first_conv": "conv_stem",
48
+ "classifier": "classifier"
49
+ }
50
+ }
preprocessor_config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architecture": "tf_efficientnetv2_s",
3
+ "data_config": {
4
+ "input_size": [
5
+ 3,
6
+ 384,
7
+ 256
8
+ ],
9
+ "interpolation": "bicubic",
10
+ "mean": [
11
+ 0.5,
12
+ 0.5,
13
+ 0.5
14
+ ],
15
+ "std": [
16
+ 0.5,
17
+ 0.5,
18
+ 0.5
19
+ ],
20
+ "resize_mode": "longest_max_size",
21
+ "padding": "center",
22
+ "pad_value": [
23
+ 127,
24
+ 127,
25
+ 127
26
+ ]
27
+ }
28
+ }