File size: 9,998 Bytes
2f97e95
 
80a2432
 
 
2f97e95
 
2ba2064
413f2bb
2f97e95
 
80a2432
413f2bb
80a2432
413f2bb
 
 
 
 
 
 
 
 
 
 
 
e8f9ded
 
80a2432
 
 
 
3a3ee47
 
 
f41b955
3a3ee47
 
 
 
 
 
 
 
 
 
 
30bfad7
3a3ee47
 
 
 
 
 
 
80a2432
29e7460
 
80a2432
413f2bb
b60fc5f
29e7460
 
 
 
 
 
413f2bb
80a2432
29e7460
 
 
 
30bfad7
29e7460
80a2432
 
 
 
413f2bb
80a2432
413f2bb
 
f41b955
 
 
413f2bb
 
 
 
 
 
 
 
 
 
 
80a2432
413f2bb
 
 
80a2432
 
413f2bb
 
 
80a2432
 
b60fc5f
 
413f2bb
80a2432
 
413f2bb
 
 
80a2432
 
 
 
 
413f2bb
 
 
 
 
 
 
 
80a2432
23a8154
 
 
 
 
 
 
 
 
 
 
 
 
413f2bb
 
7113ed1
413f2bb
 
 
 
 
 
 
80a2432
 
413f2bb
 
80a2432
 
 
413f2bb
80a2432
7113ed1
413f2bb
7113ed1
 
 
 
 
 
 
 
 
 
 
 
80a2432
 
 
413f2bb
80a2432
7113ed1
80a2432
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
---
title: README
emoji: πŸ§₯
colorFrom: indigo
colorTo: green
sdk: static
pinned: false
license: mit
short_description: SOTA fashion retrieval, measured and open.
---

<section>
  <h1>Hopit AI β€” fashion retrieval, measured properly</h1>
  <p>
    We build the <b>MODA</b> family of fashion retrieval models and benchmark them
    the way we wish everyone did: <b>full corpus, one shared harness, every
    competitor under identical protocol β€” including the cells we lose.</b>
  </p>
  <p>
    <a href="https://hopit-ai.github.io/Moda/">Full benchmark page</a>
    Β·
    <a href="https://huggingface.co/spaces/HopitAI/moda-fashion-search">Interactive demo</a>
    Β·
    <a href="https://github.com/hopit-ai/Moda">Code &amp; methodology</a>
    Β·
    <a href="https://hopit.ai">hopit.ai</a>
    Β·
    <a href="https://calendly.com/arkid_/new-meeting?back=1">Book a call</a>
  </p>
</section>

<section>
  <h2>Text-to-image β€” models under 250M parameters</h2>
  <p>The size class most deployments use. Full-corpus MAP@10; best per row bold πŸ₯‡, second πŸ₯ˆ:</p>
  <table>
    <tr><th>benchmark (corpus)</th><th><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a> 203M</th><th>MODA 203M</th><th>MODA Pro Lite 213M</th></tr>
    <tr><td>KAGL (44K)</td><td>0.2769</td><td>0.2890 πŸ₯ˆ</td><td><b>0.3185</b> πŸ₯‡</td></tr>
    <tr><td>Polyvore (94K)</td><td>0.3665</td><td>0.3726 πŸ₯ˆ</td><td><b>0.3997</b> πŸ₯‡</td></tr>
    <tr><td>Atlas (78K)</td><td>0.1826</td><td>0.1884 πŸ₯ˆ</td><td><b>0.1945</b> πŸ₯‡</td></tr>
    <tr><td>Fashion200K (202K)</td><td>0.1858 πŸ₯ˆ</td><td><b>0.1947</b> πŸ₯‡</td><td>0.1802</td></tr>
    <tr><td>DeepFashion In-Shop (53K)</td><td>0.1587 πŸ₯ˆ</td><td><b>0.1703</b> πŸ₯‡</td><td>0.1031</td></tr>
    <tr><td>DeepFashion Multimodal (43K)</td><td><b>0.0148</b> πŸ₯‡</td><td>0.0147 πŸ₯ˆ</td><td>0.0118</td></tr>
  </table>
  <p>
    Every benchmark in this class is led by a MODA-family model except one, where
    frozen FashionSigLIP keeps a 0.8% edge.
    <b><a href="https://huggingface.co/HopitAI/moda-pro-lite">MODA Pro Lite</a> owns catalog
    and title search</b> (KAGL +10.2%, Polyvore +7.3% over MODA, both significant) from one
    checkpoint, no serving recipe; <b>MODA owns captions and instance retrieval</b>
    (4 of 6 wins over FashionSigLIP significant under paired bootstrap).
  </p>
</section>

<section>
  <h2>Text-to-image β€” all systems, including larger models</h2>
  <p>
    Full-corpus MAP@10 through one shared pipeline, identical preprocessing per
    model, no gallery subsampling anywhere. Best per row in bold with πŸ₯‡, second πŸ₯ˆ.
  </p>
  <table>
    <tr><th>benchmark (corpus)</th><th><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a><br>203M</th><th>MODA<br>203M</th><th><a href="https://huggingface.co/timm/ViT-SO400M-14-SigLIP-384">SO400M</a><br>878M</th><th><a href="https://huggingface.co/srpone/zooclaw-fashionsiglip2">ZooClaw</a><br>375M</th><th>MODA Pro Lite<br>213M</th><th>MODA Pro<br>hosted</th></tr>
    <tr><td>KAGL (44K)</td><td>0.2769</td><td>0.2890</td><td><b>0.3370</b> πŸ₯‡</td><td>0.2951</td><td>0.3185</td><td>0.3263 πŸ₯ˆ</td></tr>
    <tr><td>Polyvore (94K)</td><td>0.3665</td><td>0.3726</td><td><b>0.4378</b> πŸ₯‡</td><td>0.3804</td><td>0.3997</td><td>0.4088 πŸ₯ˆ</td></tr>
    <tr><td>Atlas (78K)</td><td>0.1826</td><td>0.1884</td><td><b>0.2309</b> πŸ₯‡</td><td>0.1583</td><td>0.1945</td><td>0.2053 πŸ₯ˆ</td></tr>
    <tr><td>Fashion200K (202K)</td><td>0.1858</td><td>0.1947 πŸ₯ˆ</td><td>0.1353</td><td>0.1775</td><td>0.1802</td><td><b>0.2101</b> πŸ₯‡</td></tr>
    <tr><td>DeepFashion In-Shop (53K)</td><td>0.1587</td><td>0.1703 πŸ₯ˆ</td><td>0.1695</td><td>0.1024</td><td>0.1031</td><td><b>0.1762</b> πŸ₯‡</td></tr>
    <tr><td>DeepFashion Multimodal (43K)</td><td><b>0.0148</b> πŸ₯‡</td><td>0.0147 πŸ₯ˆ</td><td>0.0079</td><td>0.0099</td><td>0.0118</td><td>0.0144</td></tr>
  </table>
  <p>
    <b>MODA Pro: rank 1 or 2 on every row, and on 9 of 10 cells</b> once the
    H&amp;M and ZooClaw-Fashion registers are included β€” the only system with no
    bad benchmark (+6.9% mean over MODA on these six, peak +12.9%).
    <b>MODA Pro Lite</b> is the strongest open single model ≀250M on catalog
    search: KAGL <b>+10.2%</b> and Polyvore <b>+7.3%</b> over MODA (both significant), from one checkpoint with no serving machinery. Full page with ranks,
    ops costs and losses: <a href="https://hopit-ai.github.io/Moda/">benchmark page</a>.
  </p>
</section>

<section>
  <h2>Image-to-image retrieval β€” LookBench Fine R@1</h2>
  <table>
    <tr><th>model</th><th>params</th><th>Fine R@1</th></tr>
    <tr><td><b><a href="https://huggingface.co/HopitAI/moda-fashion-distilled">MODA-SigLIP-Distilled</a></b></td><td>203M</td><td><b>67.63</b> πŸ₯‡</td></tr>
    <tr><td><a href="https://serendipityoneinc.github.io/look-bench-page/">GR-Pro</a> (closed)</td><td>n/a</td><td>67.38</td></tr>
    <tr><td><a href="https://huggingface.co/TianmuLab/Tianmu-MERE">Tianmu-MERE</a></td><td>1.24B</td><td>65.99†</td></tr>
    <tr><td><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a></td><td>203M</td><td>63.84†</td></tr>
  </table>
  <p>
    †same-harness reruns; our harness reproduces Tianmu's published 66.20 within
    0.5pt. A 203M open model above a closed commercial system and a 1.24B model.
  </p>
</section>

<section>
  <h2>The MODA family</h2>
  <table>
    <tr><th>model</th><th>what it is</th><th>availability</th></tr>
    <tr>
      <td><a href="https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m"><b>MODA</b></a> (203M)</td>
      <td>zero-new-parameter serving recipe over frozen FashionSigLIP; 4/6 statistically significant full-corpus wins over its own base model</td>
      <td>open source + open weights</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/HopitAI/moda-pro-lite"><b>MODA Pro Lite</b></a> (213M)</td>
      <td>trained encoder (verified fashion-vocab build); beats MODA on catalog search as a plain bi-encoder β€” no recipe required</td>
      <td>open weights</td>
    </tr>
    <tr>
      <td><b>MODA Pro</b></td>
      <td>our hosted retrieval system β€” rank 1 or 2 on 9 of 10 benchmark cells at single-model query latency and cost</td>
      <td>closed Β· <a href="https://hopit.ai">hosted</a></td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/HopitAI/moda-fashion-distilled"><b>MODA-SigLIP-Distilled</b></a> (203M)</td>
      <td>image-to-image specialist β€” #1 open model on LookBench</td>
      <td>open weights (+ <a href="https://huggingface.co/HopitAI/moda-fashion-matryoshka">matryoshka</a>, <a href="https://huggingface.co/HopitAI/moda-fashion-distilled-512d">512d</a>, <a href="https://huggingface.co/HopitAI/moda-fashion-vision-fp16">fp16-vision</a> variants)</td>
    </tr>
  </table>
</section>

<section>
  <h2>Why trust these numbers</h2>
  <ul>
    <li><b>Full corpus only</b> β€” no subsampled galleries; screening runs are never mixed with full-corpus rows.</li>
    <li><b>One harness</b> β€” every model, ours and competitors', runs identical preprocessing and protocol.</li>
    <li><b>Losses shown</b> β€” every model card links the cells it loses; paired-bootstrap CIs where significance is claimed.</li>
    <li><b>Reproducible</b> β€” code, harness, and per-cell receipts in the <a href="https://github.com/hopit-ai/Moda">public repo</a>.</li>
  </ul>
</section>

<section>
  <h2>Which model should I use?</h2>
  <table>
    <tr><th>your workload</th><th>use</th><th>why</th></tr>
    <tr><td>catalog / title search</td><td><a href="https://huggingface.co/HopitAI/moda-pro-lite">moda-pro-lite</a></td><td>strongest ≀250M on catalog benchmarks; plain bi-encoder, any vector DB</td></tr>
    <tr><td>caption-style / exact-item search</td><td><a href="https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m">MODA</a></td><td>leads the class on Fashion200K, In-Shop, Multimodal</td></tr>
    <tr><td>best text search, zero integration</td><td><a href="https://hopit.ai">MODA Pro</a> (hosted)</td><td>rank 1–2 on 9 of 10 benchmarks, no bad benchmark</td></tr>
    <tr><td>visually similar products (image)</td><td><a href="https://huggingface.co/HopitAI/moda-fashion-distilled">moda-fashion-distilled</a></td><td>#1 open model on LookBench</td></tr>
    <tr><td>small index / edge</td><td><a href="https://huggingface.co/HopitAI/moda-fashion-matryoshka">matryoshka</a> @256d Β· <a href="https://huggingface.co/HopitAI/moda-fashion-vision-fp16">fp16</a></td><td>3Γ— smaller index at no loss Β· 186 MB vision tower</td></tr>
  </table>
  <p>All models serve on CPU. Each card links the benchmarks it loses, too.</p>
</section>

<section>
  <h2>Quick start</h2>

Text to image (plain bi-encoder, works with any vector DB):

```python
import open_clip

model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:HopitAI/moda-pro-lite")
tokenizer = open_clip.get_tokenizer("hf-hub:HopitAI/moda-pro-lite")
```

Image to image, for finding visually similar products:

```python
import open_clip

model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:HopitAI/moda-fashion-distilled")
```

MODA's full multi-view retrieval recipe:

```bash
pip install "git+https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m"
```

```python
from moda_fashionsiglip_multiview import ModaFashionSigLIP

retriever = ModaFashionSigLIP.from_pretrained()
index = retriever.build_index(image_paths, item_ids=item_ids)
results = retriever.search("red floral summer dress", index, top_k=5)[0]
```
</section>

<section>
  <h2>Use cases</h2>
  <ul>
    <li>Search a fashion catalog with a natural-language query.</li>
    <li>Find visually similar products in a fashion catalog.</li>
    <li>Match street-style looks to shoppable items.</li>
    <li>Deduplicate product images across marketplaces.</li>
    <li>Build embedding indexes for ecommerce search and recommendations.</li>
  </ul>
</section>