File size: 11,843 Bytes
fd161da
 
 
 
 
 
 
 
 
 
5e6f6db
fd161da
 
 
 
5e6f6db
fd161da
5e6f6db
 
 
 
 
fd161da
5e6f6db
 
 
 
fd161da
5e6f6db
fd161da
5e6f6db
fd161da
5e6f6db
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5d0c1d2
 
fd161da
 
5e6f6db
 
fd161da
5e6f6db
 
fd161da
5e6f6db
fd161da
5e6f6db
 
 
 
 
fd161da
5e6f6db
 
 
fd161da
5e6f6db
fd161da
5e6f6db
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fd161da
 
 
5e6f6db
 
 
 
 
 
fd161da
5e6f6db
fd161da
5e6f6db
 
 
 
fd161da
5e6f6db
fd161da
5e6f6db
 
 
 
 
 
fd161da
5e6f6db
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
---
language:
- multilingual
license: cc-by-nc-4.0
library_name: pytorch
tags:
- text-classification
- gender-detection
- nationality
- name-analysis
- character-level
- cpu
pipeline_tag: text-classification
---

# Genderize v2

Genderize v2 estimates a **binary gender** (M/F) and a **likely country** (226
ISO-3166 alpha-2 codes) from a personal name, with probabilities. It is a
character-level classifier designed for **CPU inference** β€” no tokenizer, no
vocabulary file. It supersedes `genderize v1 (cora/ultra)`, which remains
available in this repository under the git tag **`v1`**.

> **Read the limitations before using this model.** It is measurably **weaker
> than v1 in Togo, Chad and Samoa**, and it does not improve Arabic and Hebrew.
> It is intended for **aggregate statistical analyses**, not for decisions
> about individual people.

## Package contents

| File | What it is | Licence |
|---|---|---|
| `genderize_fuso_v2/fuso.json` | fusion configuration (thresholds, weights) | CC BY-NC 4.0 |
| `genderize_fuso_v2/v1/` | internal network A: the v1 `ultra` weights, configs, calibration, class map | CC BY-NC 4.0 |
| `genderize_fuso_v2/v2/` | internal network B (codepoint-level, distilled lineage), weights, configs, calibration, class map | CC BY-NC 4.0 |
| `genderize_fuso_v2/country_prior.json` | empirical country prior used by the country head | CC BY-NC 4.0 |
| `genderize_fuso_v2/dom113_rimedio.json` | configuration of the Latin-script country-trait rule (see below) | CC BY-NC 4.0 |
| `genderize_fuso_v2/SHA256SUMS` | SHA-256 of every file in the model directory | CC BY-NC 4.0 |
| `inference.py` | inference module for network A (byte-level model) | MIT |
| `inference_fuso.py` | the Genderize v2 class `GenderizeFuso` (imports `inference.py`) | MIT |
| `examples.txt` | example outputs reproduced with exactly these files | β€” |
| `LICENSE` | CC BY-NC 4.0 (weights and model artifacts) | β€” |
| `LICENSE-CODE` | MIT (inference code) | β€” |
| `COMMERCIAL_USE.md` | what counts as non-commercial use, and how to obtain a commercial licence | β€” |

Training data is **not** distributed with this package (see *Data provenance*).

## Quickstart

Requires Python 3.10+, `torch`, `numpy`, and `anyascii` (ISC licence; only
invoked for non-Latin scripts that benefit from transliteration).

```python
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("textpie/genderize")
sys.path.insert(0, path)                     # inference_fuso.py + inference.py live at repo root

from inference_fuso import GenderizeFuso     # the class documented in the file header

model = GenderizeFuso(f"{path}/genderize_fuso_v2")
print(model.predict(["Andrea Rossi", "Yuki Tanaka"]))
# -> [{'gender': 'M', 'p_male': 0.7937, 'country': 'IT', 'p_country': 0.8262, 'top5': ['IT', 'FR', 'US', 'CH', 'DE']},
#     {'gender': 'F', 'p_male': 0.5037, 'country': 'JP', 'p_country': 0.9809, 'top5': ['JP', 'US', 'ID', 'DE', 'CN']}]
```

`predict()` returns one dict per name: `gender` (`'M'` if p(M) β‰₯ 0.60, else
`'F'`), `p_male`, `country` (argmax of the adjusted country distribution),
`p_country`, and `top5` (top-5 country codes).

## How it works

- **Two internal networks, one fused output.** Network A is the v1 `ultra`
  byte-level CNN (UTF-8 bytes, 48-byte truncation). Network B reads the name
  as Unicode codepoints (32 max, its own 226-country map). Gender probability
  is the plain average of the two calibrated p(M); the country distribution is
  0.7 Γ— A + 0.3 Γ— B.
- **Calibration and threshold.** Each network ships per-head temperature
  calibration files. The operating threshold M β‰₯ **0.60** was chosen on three
  frozen benches and then verified β€” without retuning β€” on six fresh,
  never-touched sets.
- **Script-aware routing (T1).** The Unicode script of the name is detected
  from character blocks; for measured scripts the model classifies the native
  form, an `anyascii` transliteration, or the average of both (rule fixed on a
  validation half only, no learned parameters). Latin and unmapped scripts
  stay native. The country head always reads the original form.
- **Country prior.** The country distribution is adjusted with a logit term
  `log p(c) βˆ’ 0.75 Β· log Ο€(c)`, where `Ο€` is the empirical prior in
  `country_prior.json` (coefficient fixed on validation only).
- **Per-country gender selector + Latin-script trait rule.** For a small set
  of countries where v2.2 measured regressions (AM, GE, MM, TG, TD, WS), part
  of the gender decision falls back to the v1 branch under rules fixed on old
  benches and on a frozen validation bench β€” never on the 226-v2 test bench
  (configuration in `dom113_rimedio.json`).

## Results (gender accuracy, measured)

v1 = `genderize-ultra` (calibrated, threshold 0.50). v2 = this release
(threshold 0.60). All numbers are point measurements on frozen benches; they
do not generalize beyond the named benches.

### Independent bench: 226-countries v2 (50,258 rows, Latin script)

Built with explicit exclusions against the training corpora and all earlier
benches; **no overlap with training** (full = residual by construction).
Paired bootstrap, 2,000 resamples, seed 20260927.

| Group | n | v2 | v1 |
|---|---:|---:|---:|
| All | 50,258 | **94.08 %** | 93.71 % |
| Women (F) | 20,648 | **92.98 %** | 90.30 % |
| Men (M) | 29,610 | 94.86 % | **96.09 %** |

Per country (countries with β‰₯ 400 bench rows): **14 countries significantly
above v1, 94 statistically par, 1 below**. With the pre-registered criterion
(β‰₯ 200 rows): **14 above, 106 par, 3 below**. Countries above v1: AZ, DZ,
EG, ET, HK, JM, LK, MY, NG, NP, SG, SY, UY, ZA. "Par" means the confidence
interval includes zero β€” it is not proof of equivalence; no multiple-comparison
correction is applied.

### Where v2 is **worse than v1** (all measured deltas, IC95 in percentage points)

| Country | n | Ξ” (v2 βˆ’ v1) | IC95 |
|---|---:|---:|---|
| **Togo (TG)** | 400 | **βˆ’2.00** | [βˆ’3.50, βˆ’0.50] |
| **Chad (TD)** | 207 | **βˆ’2.42** | [βˆ’4.83, βˆ’0.48] |
| **Samoa (WS)** | 210 | **βˆ’4.76** | [βˆ’9.05, βˆ’0.95] |
| Armenia (AM) | 400 | βˆ’0.25 | [βˆ’0.75, 0.00] |
| Georgia (GE) | 400 | βˆ’0.25 | [βˆ’0.75, 0.00] |
| Myanmar (MM) | 400 | βˆ’1.25 | [βˆ’3.50, +0.75] |

TG is the only country below v1 at the β‰₯ 400-row quota; TD and WS are also
below at the pre-registered β‰₯ 200-row criterion; AM, GE and MM measure
statistically par. Without the Latin-script trait rule, deltas down to
βˆ’7.3 pp were measured on the same bench. **If your
analysis focuses on these countries, use v1.**

### Other frozen benches

| Bench | n | v2 | v1 | Notes |
|---|---:|---:|---:|---|
| Native scripts (with T1 routing) | 16,945 | **88.70 %** | 83.84 % | release path (v2.2 verdict) |
| Native scripts + routing (T1) | 6,072 | **85.82 %** | 74.31 % | half-test split; largest gains: Georgian 50.9β†’92.3, Bengali 62.5β†’87.7, Armenian 67.5β†’90.8, Kazakh 93.1β†’97.5, Hindi 83.8β†’87.3 |
| East-Asian | 5,199 | **80.96 %** | 79.75 % | women 79.3 vs 69.3 |
| Old 226 bench (Wikidata) | 39,164 | 92.63 % | 92.14 % | **not independent: 34 % of this bench is inside v1's training data** (v1 residual: 90.41 %); kept for continuity with v1's history |

Six fresh, never-touched sets (v2 / v1): Olympians 96.20 / 96.07; Vietnamese
80.32 / 76.58; Tajik-Latin 79.09 / 76.13; Tajik-Cyrillic 72.12 / 67.65;
Persian-Latin 77.57 / 74.07; Persian-script 67.29 / 64.59.

### Country (nationality) output

The country head is a **coarse signal**: on the old 226 bench, top-1 β‰ˆ 28.4 %
(v1: 28.1 %; v1 top-5 β‰ˆ 56.3 %). Use `top5` and the probability, not the
argmax alone.

### Persian dedicated branch: tested and **rejected**

A dedicated Persian transliteration branch (beam over k readings, < 5 MB) was
integrated and measured on the only available native-Persian bench (n = 1,000;
**not fresh** β€” already used by earlier experiments): +2.5 pp over the direct
branch, IC95 [βˆ’0.4, +5.6] β€” the interval crosses zero, so the gain is **not
confirmed** and the branch is **not part of this release**. It will be
reconsidered only if an independent, fresh native-Persian bench becomes
available (none found as of 2026-09-28).

### Speed

198.8 names/s with script routing (189.6 without), single CPU, 8 threads,
batch 256, measured back-to-back on the bench machine. Throughput varies with
hardware and load; no latency promise is made.

## Limitations and honest notes

- **Weaker than v1 in Togo, Chad, Samoa** (deltas above) β€” the headline
  limitation of this release. Armenia, Georgia and Myanmar measure slightly
  below v1 (statistically par).
- **Women with non-Latin-script names remain weak**: on Persian-script and
  Cyrillic names female accuracy stays around 29–35 %, as low as 19.9 % on
  Persian-script names in the fresh sets. The aggregate gains do not fix this.
- **Arabic and Hebrew: no gain.** Transliteration does not help these scripts
  (measured: routing leaves them at the native branch); v2 is at best par with
  v1 there. Thai likewise stays native.
- **The model does not condition gender on country**: e.g. "Andrea Costa" is
  classified F with p(M) = 0.44 (Andrea is a male name in Italian, female in
  Spanish).
- **No built-in abstention**: the model always answers. Probabilities are
  temperature-calibrated per network; the 0.60 threshold was fixed before the
  fresh-set verification. For analyses, use `p_male` (and a margin you choose)
  rather than the hard label.
- **Binary output only**: the model returns M/F. It cannot represent
  non-binary identities, and name-inferred gender is a *perceived-attribute
  proxy*, not a statement about any person's identity.
- Bench numbers are specific to the named frozen benches; several are
  Wikidata-derived and may share source characteristics with the training
  corpora (the independent 226-v2 bench was built with explicit exclusions;
  the old 226 bench overlaps v1 training by 34 %, as declared above).

## Data provenance

- Trained on **aggregated public name records with gender and country signals**
  (Wikidata-derived corpora, Olympian records, per-country name lists),
  internally reviewed. Training corpora, splits and recipes are **private and
  not redistributed** with this package.
- **No LLM-generated labels were used in the training of this model.** An
  LLM second-opinion cascade was explored in research and measured, but it is
  **not part of this release**.
- A small fraction of the network-B corpus (**8,708 rows, β‰ˆ 1.21 %**) comes
  from a GFDL-1.3-licensed public dataset; this is disclosed for transparency
  about obligations.
- WGND and IPUMS were never used in training.
- Bench sources: Wikidata (CC0), Olympian and public record compilations;
  bench row-level data is not redistributed.

## Licence

- **Weights and model artifacts** (everything under `genderize_fuso_v2/`):
  **CC BY-NC 4.0** β€” see `LICENSE`; commercial use requires a separate
  licence (see `COMMERCIAL_USE.md`).
- **Inference code** (`inference.py`, `inference_fuso.py`): **MIT** β€” see
  `LICENSE-CODE`.
- Runtime dependency `anyascii`: ISC licence (third-party).

## Intended use

Aggregate, statistical analyses: bibliometrics and authorship studies,
name-collection demography, dataset documentation. **Not** intended for
decisions about individual people (hiring, credit, identity verification,
content moderation), nor as ground truth about anyone's gender.

## Version history

- **v2** (this release): fused two-network model, script-aware routing,
  country prior, per-country selector. 14 above, 94 par, 1 below at β‰₯ 400 rows
  (3 below at β‰₯ 200) on the independent 226-v2 bench (see above).
- **v1** (`cora` / `ultra`): preserved in this repository at git tag **`v1`**.

---

Support this work: https://github.com/sponsors/sheppard94g