File size: 2,043 Bytes
5576acd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---

language:
  - en
  - zh
  - de
  - es
  - fr
  - ja
library_name: pytorch
pipeline_tag: text-classification
base_model: google/embeddinggemma-300m
tags:
  - extreme-multi-label-classification
  - embeddinggemma
  - multilingual
---


# UserNeeds1K (FP16)

This standalone variant stores all weights in FP16. EmbeddingGemma computes in BF16 for numerical stability, while the classifier computes in FP16.

An extreme multi-label classifier for 1,091 categories. Text is embedded with
the bundled `google_embeddinggemma-300m` model; a dynamic-negative MLP produces
independent category scores. The package runs without downloading a separate
embedding model.

## Usage

```bash

pip install -r requirements.txt

python example.py

```

Example output:

```text

0.6411  10004/20035/30203  Beauty & Fitness / Fitness / Other

0.4917  10004/20035/30200  Beauty & Fitness / Fitness / Fitness Instruction & Personal Training

0.2998  10022/20214/30775  Reference / General Reference / How-To, DIY & Expert Content

0.0853  10004/20035/30202  Beauty & Fitness / Fitness / High Intensity Interval Training

0.0628  10002/20013/30078  Arts & Entertainment / TV & Video / Online Video

```

## Supported language keys

| Key | Language |
|---|---|
| `cn` | Chinese |
| `en` | English |
| `de` | German |
| `ge` | German alias |
| `sp` | Spanish |
| `fr` | French |
| `jp` | Japanese |

## Performance

| Test set | nDCG@5 | Precision@5 | Recall@5 | Top-1 hit | Tail recall@5 |
|---|---:|---:|---:|---:|---:|
| English | 0.858722 | 0.513768 | 0.833800 | 0.910088 | 0.629620 |
| Chinese | 0.821382 | 0.491000 | 0.792756 | 0.890000 | 0.583333 |

On an H100, this optimized FP16-storage variant processed 825.51 examples per
second with 2,176.64 MB peak GPU memory and achieved 0.855899 nDCG@5 on the
1,000-example precision comparison set.

No training records, translated samples, document embeddings, targets, split
files, or training-derived document centroids are included in this package.