| --- |
| title: README |
| emoji: π§₯ |
| colorFrom: indigo |
| colorTo: green |
| sdk: static |
| pinned: false |
| license: mit |
| short_description: SOTA fashion retrieval, measured and open. |
| --- |
| |
| <section> |
| <h1>Hopit AI β fashion retrieval, measured properly</h1> |
| <p> |
| We build the <b>MODA</b> family of fashion retrieval models and benchmark them |
| the way we wish everyone did: <b>full corpus, one shared harness, every |
| competitor under identical protocol β including the cells we lose.</b> |
| </p> |
| <p> |
| <a href="https://hopit-ai.github.io/Moda/">Full benchmark page</a> |
| Β· |
| <a href="https://huggingface.co/spaces/HopitAI/moda-fashion-search">Interactive demo</a> |
| Β· |
| <a href="https://github.com/hopit-ai/Moda">Code & methodology</a> |
| Β· |
| <a href="https://hopit.ai">hopit.ai</a> |
| Β· |
| <a href="https://calendly.com/arkid_/new-meeting?back=1">Book a call</a> |
| </p> |
| </section> |
| |
| <section> |
| <h2>Text-to-image β models under 250M parameters</h2> |
| <p>The size class most deployments use. Full-corpus MAP@10; best per row bold π₯, second π₯:</p> |
| <table> |
| <tr><th>benchmark (corpus)</th><th><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a> 203M</th><th>MODA 203M</th><th>MODA Pro Lite 213M</th></tr> |
| <tr><td>KAGL (44K)</td><td>0.2769</td><td>0.2890 π₯</td><td><b>0.3185</b> π₯</td></tr> |
| <tr><td>Polyvore (94K)</td><td>0.3665</td><td>0.3726 π₯</td><td><b>0.3997</b> π₯</td></tr> |
| <tr><td>Atlas (78K)</td><td>0.1826</td><td>0.1884 π₯</td><td><b>0.1945</b> π₯</td></tr> |
| <tr><td>Fashion200K (202K)</td><td>0.1858 π₯</td><td><b>0.1947</b> π₯</td><td>0.1802</td></tr> |
| <tr><td>DeepFashion In-Shop (53K)</td><td>0.1587 π₯</td><td><b>0.1703</b> π₯</td><td>0.1031</td></tr> |
| <tr><td>DeepFashion Multimodal (43K)</td><td><b>0.0148</b> π₯</td><td>0.0147 π₯</td><td>0.0118</td></tr> |
| </table> |
| <p> |
| Every benchmark in this class is led by a MODA-family model except one, where |
| frozen FashionSigLIP keeps a 0.8% edge. |
| <b><a href="https://huggingface.co/HopitAI/moda-pro-lite">MODA Pro Lite</a> owns catalog |
| and title search</b> (KAGL +10.2%, Polyvore +7.3% over MODA, both significant) from one |
| checkpoint, no serving recipe; <b>MODA owns captions and instance retrieval</b> |
| (4 of 6 wins over FashionSigLIP significant under paired bootstrap). |
| </p> |
| </section> |
| |
| <section> |
| <h2>Text-to-image β all systems, including larger models</h2> |
| <p> |
| Full-corpus MAP@10 through one shared pipeline, identical preprocessing per |
| model, no gallery subsampling anywhere. Best per row in bold with π₯, second π₯. |
| </p> |
| <table> |
| <tr><th>benchmark (corpus)</th><th><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a><br>203M</th><th>MODA<br>203M</th><th><a href="https://huggingface.co/timm/ViT-SO400M-14-SigLIP-384">SO400M</a><br>878M</th><th><a href="https://huggingface.co/srpone/zooclaw-fashionsiglip2">ZooClaw</a><br>375M</th><th>MODA Pro Lite<br>213M</th><th>MODA Pro<br>hosted</th></tr> |
| <tr><td>KAGL (44K)</td><td>0.2769</td><td>0.2890</td><td><b>0.3370</b> π₯</td><td>0.2951</td><td>0.3185</td><td>0.3263 π₯</td></tr> |
| <tr><td>Polyvore (94K)</td><td>0.3665</td><td>0.3726</td><td><b>0.4378</b> π₯</td><td>0.3804</td><td>0.3997</td><td>0.4088 π₯</td></tr> |
| <tr><td>Atlas (78K)</td><td>0.1826</td><td>0.1884</td><td><b>0.2309</b> π₯</td><td>0.1583</td><td>0.1945</td><td>0.2053 π₯</td></tr> |
| <tr><td>Fashion200K (202K)</td><td>0.1858</td><td>0.1947 π₯</td><td>0.1353</td><td>0.1775</td><td>0.1802</td><td><b>0.2101</b> π₯</td></tr> |
| <tr><td>DeepFashion In-Shop (53K)</td><td>0.1587</td><td>0.1703 π₯</td><td>0.1695</td><td>0.1024</td><td>0.1031</td><td><b>0.1762</b> π₯</td></tr> |
| <tr><td>DeepFashion Multimodal (43K)</td><td><b>0.0148</b> π₯</td><td>0.0147 π₯</td><td>0.0079</td><td>0.0099</td><td>0.0118</td><td>0.0144</td></tr> |
| </table> |
| <p> |
| <b>MODA Pro: rank 1 or 2 on every row, and on 9 of 10 cells</b> once the |
| H&M and ZooClaw-Fashion registers are included β the only system with no |
| bad benchmark (+6.9% mean over MODA on these six, peak +12.9%). |
| <b>MODA Pro Lite</b> is the strongest open single model β€250M on catalog |
| search: KAGL <b>+10.2%</b> and Polyvore <b>+7.3%</b> over MODA (both significant), from one checkpoint with no serving machinery. Full page with ranks, |
| ops costs and losses: <a href="https://hopit-ai.github.io/Moda/">benchmark page</a>. |
| </p> |
| </section> |
| |
| <section> |
| <h2>Image-to-image retrieval β LookBench Fine R@1</h2> |
| <table> |
| <tr><th>model</th><th>params</th><th>Fine R@1</th></tr> |
| <tr><td><b><a href="https://huggingface.co/HopitAI/moda-fashion-distilled">MODA-SigLIP-Distilled</a></b></td><td>203M</td><td><b>67.63</b> π₯</td></tr> |
| <tr><td><a href="https://serendipityoneinc.github.io/look-bench-page/">GR-Pro</a> (closed)</td><td>n/a</td><td>67.38</td></tr> |
| <tr><td><a href="https://huggingface.co/TianmuLab/Tianmu-MERE">Tianmu-MERE</a></td><td>1.24B</td><td>65.99β </td></tr> |
| <tr><td><a href="https://huggingface.co/Marqo/marqo-fashionSigLIP">FashionSigLIP</a></td><td>203M</td><td>63.84β </td></tr> |
| </table> |
| <p> |
| β same-harness reruns; our harness reproduces Tianmu's published 66.20 within |
| 0.5pt. A 203M open model above a closed commercial system and a 1.24B model. |
| </p> |
| </section> |
| |
| <section> |
| <h2>The MODA family</h2> |
| <table> |
| <tr><th>model</th><th>what it is</th><th>availability</th></tr> |
| <tr> |
| <td><a href="https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m"><b>MODA</b></a> (203M)</td> |
| <td>zero-new-parameter serving recipe over frozen FashionSigLIP; 4/6 statistically significant full-corpus wins over its own base model</td> |
| <td>open source + open weights</td> |
| </tr> |
| <tr> |
| <td><a href="https://huggingface.co/HopitAI/moda-pro-lite"><b>MODA Pro Lite</b></a> (213M)</td> |
| <td>trained encoder (verified fashion-vocab build); beats MODA on catalog search as a plain bi-encoder β no recipe required</td> |
| <td>open weights</td> |
| </tr> |
| <tr> |
| <td><b>MODA Pro</b></td> |
| <td>our hosted retrieval system β rank 1 or 2 on 9 of 10 benchmark cells at single-model query latency and cost</td> |
| <td>closed Β· <a href="https://hopit.ai">hosted</a></td> |
| </tr> |
| <tr> |
| <td><a href="https://huggingface.co/HopitAI/moda-fashion-distilled"><b>MODA-SigLIP-Distilled</b></a> (203M)</td> |
| <td>image-to-image specialist β #1 open model on LookBench</td> |
| <td>open weights (+ <a href="https://huggingface.co/HopitAI/moda-fashion-matryoshka">matryoshka</a>, <a href="https://huggingface.co/HopitAI/moda-fashion-distilled-512d">512d</a>, <a href="https://huggingface.co/HopitAI/moda-fashion-vision-fp16">fp16-vision</a> variants)</td> |
| </tr> |
| </table> |
| </section> |
| |
| <section> |
| <h2>Why trust these numbers</h2> |
| <ul> |
| <li><b>Full corpus only</b> β no subsampled galleries; screening runs are never mixed with full-corpus rows.</li> |
| <li><b>One harness</b> β every model, ours and competitors', runs identical preprocessing and protocol.</li> |
| <li><b>Losses shown</b> β every model card links the cells it loses; paired-bootstrap CIs where significance is claimed.</li> |
| <li><b>Reproducible</b> β code, harness, and per-cell receipts in the <a href="https://github.com/hopit-ai/Moda">public repo</a>.</li> |
| </ul> |
| </section> |
| |
| <section> |
| <h2>Which model should I use?</h2> |
| <table> |
| <tr><th>your workload</th><th>use</th><th>why</th></tr> |
| <tr><td>catalog / title search</td><td><a href="https://huggingface.co/HopitAI/moda-pro-lite">moda-pro-lite</a></td><td>strongest β€250M on catalog benchmarks; plain bi-encoder, any vector DB</td></tr> |
| <tr><td>caption-style / exact-item search</td><td><a href="https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m">MODA</a></td><td>leads the class on Fashion200K, In-Shop, Multimodal</td></tr> |
| <tr><td>best text search, zero integration</td><td><a href="https://hopit.ai">MODA Pro</a> (hosted)</td><td>rank 1β2 on 9 of 10 benchmarks, no bad benchmark</td></tr> |
| <tr><td>visually similar products (image)</td><td><a href="https://huggingface.co/HopitAI/moda-fashion-distilled">moda-fashion-distilled</a></td><td>#1 open model on LookBench</td></tr> |
| <tr><td>small index / edge</td><td><a href="https://huggingface.co/HopitAI/moda-fashion-matryoshka">matryoshka</a> @256d Β· <a href="https://huggingface.co/HopitAI/moda-fashion-vision-fp16">fp16</a></td><td>3Γ smaller index at no loss Β· 186 MB vision tower</td></tr> |
| </table> |
| <p>All models serve on CPU. Each card links the benchmarks it loses, too.</p> |
| </section> |
| |
| <section> |
| <h2>Quick start</h2> |
|
|
| Text to image (plain bi-encoder, works with any vector DB): |
|
|
| ```python |
| import open_clip |
| |
| model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:HopitAI/moda-pro-lite") |
| tokenizer = open_clip.get_tokenizer("hf-hub:HopitAI/moda-pro-lite") |
| ``` |
|
|
| Image to image, for finding visually similar products: |
|
|
| ```python |
| import open_clip |
| |
| model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:HopitAI/moda-fashion-distilled") |
| ``` |
|
|
| MODA's full multi-view retrieval recipe: |
|
|
| ```bash |
| pip install "git+https://huggingface.co/HopitAI/moda-fashionsiglip-multiview-203m" |
| ``` |
|
|
| ```python |
| from moda_fashionsiglip_multiview import ModaFashionSigLIP |
| |
| retriever = ModaFashionSigLIP.from_pretrained() |
| index = retriever.build_index(image_paths, item_ids=item_ids) |
| results = retriever.search("red floral summer dress", index, top_k=5)[0] |
| ``` |
| </section> |
|
|
| <section> |
| <h2>Use cases</h2> |
| <ul> |
| <li>Search a fashion catalog with a natural-language query.</li> |
| <li>Find visually similar products in a fashion catalog.</li> |
| <li>Match street-style looks to shoppable items.</li> |
| <li>Deduplicate product images across marketplaces.</li> |
| <li>Build embedding indexes for ecommerce search and recommendations.</li> |
| </ul> |
| </section> |
| |