File size: 5,833 Bytes
cd1b764
 
6fd27bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cd1b764
6fd27bf
 
 
669d050
6fd27bf
 
 
 
 
 
 
669d050
 
 
6fd27bf
 
 
 
 
 
 
 
 
 
669d050
6fd27bf
669d050
6fd27bf
 
 
 
669d050
6fd27bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b69eefc
 
6fd27bf
 
 
 
 
 
b69eefc
6fd27bf
 
669d050
6fd27bf
 
 
 
 
 
b69eefc
 
6fd27bf
 
b69eefc
 
 
 
6fd27bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
669d050
6fd27bf
 
 
 
 
 
 
669d050
6fd27bf
669d050
 
 
 
 
 
 
6fd27bf
 
 
 
 
669d050
 
 
 
 
 
 
6fd27bf
 
 
 
 
 
 
 
669d050
6fd27bf
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
---
license: mit
library_name: timm
pipeline_tag: image-classification
base_model: timm/tf_efficientnetv2_s.in21k_ft_in1k
base_model_relation: finetune
model_name: Clothing Type EfficientNetV2-S
metrics:
  - accuracy
tags:
  - image-classification
  - timm
  - pytorch
  - efficientnetv2
  - onnx
  - safetensors
  - clothing-type-classification
---

# Clothing Type EfficientNetV2-S

This is a five-class clothing-type image classifier fine-tuned from
[`timm/tf_efficientnetv2_s.in21k_ft_in1k`](https://huggingface.co/timm/tf_efficientnetv2_s.in21k_ft_in1k).
It predicts one of the following labels, in model-output order:

| ID | Label |
|---:|---|
| 0 | `denim_shorts` |
| 1 | `lingerie` |
| 2 | `maid` |
| 3 | `nurse` |
| 4 | `swim_bikini` |

## Model details

| Property | Value |
|---|---|
| Task | Image classification |
| Architecture | EfficientNetV2-S |
| Base model | `timm/tf_efficientnetv2_s.in21k_ft_in1k` |
| Relationship | Full fine-tune |
| Precision | FP32; not quantized |
| Parameters | 20,183,893 |
| Input | RGB image tensor, `N × 3 × 384 × 256` |
| Output | Five logits in the label order above |
| ONNX opset | 18 |
| License | MIT; the base model is Apache-2.0 |

The pretrained 1,000-class head is replaced by a dropout layer with
probability `0.15` and a five-output linear classifier. The published
artifacts are full-precision exports, not quantized variants. The ONNX model
supports dynamic batch sizes and fixes each image to 384 pixels high by 256
pixels wide.

## Files

- `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx` is the portable
  inference graph.
- `clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors` contains the
  PyTorch `EfficientNetClassifier` state dictionary.
- `config.json` records the architecture, labels, and timm model settings.
- `preprocessor_config.json` records the image preprocessing contract.

The safetensors keys include the custom wrapper's `model.` prefix and custom
classifier head. Load them with the `EfficientNetClassifier` implementation
from the [source repository](https://github.com/mr-szgz/clothing-classification-model),
not as an unmodified upstream timm checkpoint.

## Preprocessing

For the model tensor itself:

1. Convert the image to RGB.
2. Resize it with preserved aspect ratio to fit within `256 × 384`, using
   bicubic interpolation when enlarging and area interpolation when reducing.
3. Center-pad it to `256 × 384` with RGB value `(127, 127, 127)`.
4. Rescale unsigned 8-bit pixels by `1 / 255`.
5. Normalize each channel with mean `0.5` and standard deviation `0.5`.
6. Convert from HWC to NCHW layout and add the batch dimension.

The source data pipeline first detects the largest person with YOLO and crops
to that bounding box. Inputs that are already tightly person-centered are
closest to the training distribution. Person detection is not part of the
published classifier and is unnecessary when the input is already cropped.

## ONNX inference

Install `huggingface_hub`, `transformers`, `timm`, `torchvision`, `onnxruntime`,
`numpy`, and `Pillow`, then run:

```python
from huggingface_hub import hf_hub_download
import numpy as np
import onnxruntime as ort
from PIL import Image, ImageOps
from transformers import AutoImageProcessor

REPO_ID = "electblake/clothing_type_classifier"
LABELS = ["denim_shorts", "lingerie", "maid", "nurse", "swim_bikini"]
model_path = hf_hub_download(
    REPO_ID,
    "clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.onnx",
)

image = ImageOps.exif_transpose(Image.open("person.jpg")).convert("RGB")
image_processor = AutoImageProcessor.from_pretrained(REPO_ID)
inputs = image_processor(image, return_tensors="pt")

session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
logits = session.run(
    ["logits"],
    {"image": inputs["pixel_values"].numpy()},
)[0][0]
probabilities = np.exp(logits - logits.max())
probabilities /= probabilities.sum()

prediction = int(probabilities.argmax())
print(LABELS[prediction], float(probabilities[prediction]))
```

## Safetensors inference

From a checkout of the source repository:

```python
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

from src.model import EfficientNetClassifier

weights_path = hf_hub_download(
    "electblake/clothing_type_classifier",
    "clothing_type-tf_efficientnetv2_s.in21k_ft_in1k.safetensors",
)
model = EfficientNetClassifier(
    model_name="tf_efficientnetv2_s.in21k_ft_in1k",
    num_classes=5,
    dropout=0.15,
    pretrained=False,
)
model.load_state_dict(load_file(weights_path))
model.eval()
```

## Training data

The prepared dataset contains 16,496 person crops across the five labels.

| Split | Images |
|---|---:|
| Train | 11,545 |
| Validation | 2,474 |
| Test | 2,477 |

Training used weighted cross-entropy, AdamW, MixUp, CutMix, horizontal flips,
geometric transforms, brightness and contrast changes, and hue and saturation
changes. The backbone was initially frozen and then fully fine-tuned.

## Evaluation

The published checkpoint was selected at epoch 31 with **89.73% validation
accuracy**. This value comes from the checkpoint metadata and is a
model-selection result on the project's validation split, not an independent
benchmark or test-set result.

## Intended use and limitations

The model is intended for research, media organization, and non-critical
clothing-type tagging. It is not intended for decisions affecting a person's
rights or access to services.

Performance is expected to vary with occlusion, multiple people, unusual
framing, small subjects, lighting, image quality, and clothing outside the
five published categories. The classifier always selects one of its five
known labels and has no `other` or rejection class. Confidence scores are
softmax probabilities and have not been calibrated.