genderize / README.md
textpie's picture
Card: declare numpy runtime dependency (caught by clean-environment verification)
5d0c1d2 verified
|
Raw History Blame Contribute Delete
11.8 kB
---
language:
- multilingual
license: cc-by-nc-4.0
library_name: pytorch
tags:
- text-classification
- gender-detection
- nationality
- name-analysis
- character-level
- cpu
pipeline_tag: text-classification
---
# Genderize v2
Genderize v2 estimates a **binary gender** (M/F) and a **likely country** (226
ISO-3166 alpha-2 codes) from a personal name, with probabilities. It is a
character-level classifier designed for **CPU inference** β€” no tokenizer, no
vocabulary file. It supersedes `genderize v1 (cora/ultra)`, which remains
available in this repository under the git tag **`v1`**.
> **Read the limitations before using this model.** It is measurably **weaker
> than v1 in Togo, Chad and Samoa**, and it does not improve Arabic and Hebrew.
> It is intended for **aggregate statistical analyses**, not for decisions
> about individual people.
## Package contents
| File | What it is | Licence |
|---|---|---|
| `genderize_fuso_v2/fuso.json` | fusion configuration (thresholds, weights) | CC BY-NC 4.0 |
| `genderize_fuso_v2/v1/` | internal network A: the v1 `ultra` weights, configs, calibration, class map | CC BY-NC 4.0 |
| `genderize_fuso_v2/v2/` | internal network B (codepoint-level, distilled lineage), weights, configs, calibration, class map | CC BY-NC 4.0 |
| `genderize_fuso_v2/country_prior.json` | empirical country prior used by the country head | CC BY-NC 4.0 |
| `genderize_fuso_v2/dom113_rimedio.json` | configuration of the Latin-script country-trait rule (see below) | CC BY-NC 4.0 |
| `genderize_fuso_v2/SHA256SUMS` | SHA-256 of every file in the model directory | CC BY-NC 4.0 |
| `inference.py` | inference module for network A (byte-level model) | MIT |
| `inference_fuso.py` | the Genderize v2 class `GenderizeFuso` (imports `inference.py`) | MIT |
| `examples.txt` | example outputs reproduced with exactly these files | β€” |
| `LICENSE` | CC BY-NC 4.0 (weights and model artifacts) | β€” |
| `LICENSE-CODE` | MIT (inference code) | β€” |
| `COMMERCIAL_USE.md` | what counts as non-commercial use, and how to obtain a commercial licence | β€” |
Training data is **not** distributed with this package (see *Data provenance*).
## Quickstart
Requires Python 3.10+, `torch`, `numpy`, and `anyascii` (ISC licence; only
invoked for non-Latin scripts that benefit from transliteration).
```python
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("textpie/genderize")
sys.path.insert(0, path) # inference_fuso.py + inference.py live at repo root
from inference_fuso import GenderizeFuso # the class documented in the file header
model = GenderizeFuso(f"{path}/genderize_fuso_v2")
print(model.predict(["Andrea Rossi", "Yuki Tanaka"]))
# -> [{'gender': 'M', 'p_male': 0.7937, 'country': 'IT', 'p_country': 0.8262, 'top5': ['IT', 'FR', 'US', 'CH', 'DE']},
# {'gender': 'F', 'p_male': 0.5037, 'country': 'JP', 'p_country': 0.9809, 'top5': ['JP', 'US', 'ID', 'DE', 'CN']}]
```
`predict()` returns one dict per name: `gender` (`'M'` if p(M) β‰₯ 0.60, else
`'F'`), `p_male`, `country` (argmax of the adjusted country distribution),
`p_country`, and `top5` (top-5 country codes).
## How it works
- **Two internal networks, one fused output.** Network A is the v1 `ultra`
byte-level CNN (UTF-8 bytes, 48-byte truncation). Network B reads the name
as Unicode codepoints (32 max, its own 226-country map). Gender probability
is the plain average of the two calibrated p(M); the country distribution is
0.7 Γ— A + 0.3 Γ— B.
- **Calibration and threshold.** Each network ships per-head temperature
calibration files. The operating threshold M β‰₯ **0.60** was chosen on three
frozen benches and then verified β€” without retuning β€” on six fresh,
never-touched sets.
- **Script-aware routing (T1).** The Unicode script of the name is detected
from character blocks; for measured scripts the model classifies the native
form, an `anyascii` transliteration, or the average of both (rule fixed on a
validation half only, no learned parameters). Latin and unmapped scripts
stay native. The country head always reads the original form.
- **Country prior.** The country distribution is adjusted with a logit term
`log p(c) βˆ’ 0.75 Β· log Ο€(c)`, where `Ο€` is the empirical prior in
`country_prior.json` (coefficient fixed on validation only).
- **Per-country gender selector + Latin-script trait rule.** For a small set
of countries where v2.2 measured regressions (AM, GE, MM, TG, TD, WS), part
of the gender decision falls back to the v1 branch under rules fixed on old
benches and on a frozen validation bench β€” never on the 226-v2 test bench
(configuration in `dom113_rimedio.json`).
## Results (gender accuracy, measured)
v1 = `genderize-ultra` (calibrated, threshold 0.50). v2 = this release
(threshold 0.60). All numbers are point measurements on frozen benches; they
do not generalize beyond the named benches.
### Independent bench: 226-countries v2 (50,258 rows, Latin script)
Built with explicit exclusions against the training corpora and all earlier
benches; **no overlap with training** (full = residual by construction).
Paired bootstrap, 2,000 resamples, seed 20260927.
| Group | n | v2 | v1 |
|---|---:|---:|---:|
| All | 50,258 | **94.08 %** | 93.71 % |
| Women (F) | 20,648 | **92.98 %** | 90.30 % |
| Men (M) | 29,610 | 94.86 % | **96.09 %** |
Per country (countries with β‰₯ 400 bench rows): **14 countries significantly
above v1, 94 statistically par, 1 below**. With the pre-registered criterion
(β‰₯ 200 rows): **14 above, 106 par, 3 below**. Countries above v1: AZ, DZ,
EG, ET, HK, JM, LK, MY, NG, NP, SG, SY, UY, ZA. "Par" means the confidence
interval includes zero β€” it is not proof of equivalence; no multiple-comparison
correction is applied.
### Where v2 is **worse than v1** (all measured deltas, IC95 in percentage points)
| Country | n | Ξ” (v2 βˆ’ v1) | IC95 |
|---|---:|---:|---|
| **Togo (TG)** | 400 | **βˆ’2.00** | [βˆ’3.50, βˆ’0.50] |
| **Chad (TD)** | 207 | **βˆ’2.42** | [βˆ’4.83, βˆ’0.48] |
| **Samoa (WS)** | 210 | **βˆ’4.76** | [βˆ’9.05, βˆ’0.95] |
| Armenia (AM) | 400 | βˆ’0.25 | [βˆ’0.75, 0.00] |
| Georgia (GE) | 400 | βˆ’0.25 | [βˆ’0.75, 0.00] |
| Myanmar (MM) | 400 | βˆ’1.25 | [βˆ’3.50, +0.75] |
TG is the only country below v1 at the β‰₯ 400-row quota; TD and WS are also
below at the pre-registered β‰₯ 200-row criterion; AM, GE and MM measure
statistically par. Without the Latin-script trait rule, deltas down to
βˆ’7.3 pp were measured on the same bench. **If your
analysis focuses on these countries, use v1.**
### Other frozen benches
| Bench | n | v2 | v1 | Notes |
|---|---:|---:|---:|---|
| Native scripts (with T1 routing) | 16,945 | **88.70 %** | 83.84 % | release path (v2.2 verdict) |
| Native scripts + routing (T1) | 6,072 | **85.82 %** | 74.31 % | half-test split; largest gains: Georgian 50.9β†’92.3, Bengali 62.5β†’87.7, Armenian 67.5β†’90.8, Kazakh 93.1β†’97.5, Hindi 83.8β†’87.3 |
| East-Asian | 5,199 | **80.96 %** | 79.75 % | women 79.3 vs 69.3 |
| Old 226 bench (Wikidata) | 39,164 | 92.63 % | 92.14 % | **not independent: 34 % of this bench is inside v1's training data** (v1 residual: 90.41 %); kept for continuity with v1's history |
Six fresh, never-touched sets (v2 / v1): Olympians 96.20 / 96.07; Vietnamese
80.32 / 76.58; Tajik-Latin 79.09 / 76.13; Tajik-Cyrillic 72.12 / 67.65;
Persian-Latin 77.57 / 74.07; Persian-script 67.29 / 64.59.
### Country (nationality) output
The country head is a **coarse signal**: on the old 226 bench, top-1 β‰ˆ 28.4 %
(v1: 28.1 %; v1 top-5 β‰ˆ 56.3 %). Use `top5` and the probability, not the
argmax alone.
### Persian dedicated branch: tested and **rejected**
A dedicated Persian transliteration branch (beam over k readings, < 5 MB) was
integrated and measured on the only available native-Persian bench (n = 1,000;
**not fresh** β€” already used by earlier experiments): +2.5 pp over the direct
branch, IC95 [βˆ’0.4, +5.6] β€” the interval crosses zero, so the gain is **not
confirmed** and the branch is **not part of this release**. It will be
reconsidered only if an independent, fresh native-Persian bench becomes
available (none found as of 2026-09-28).
### Speed
198.8 names/s with script routing (189.6 without), single CPU, 8 threads,
batch 256, measured back-to-back on the bench machine. Throughput varies with
hardware and load; no latency promise is made.
## Limitations and honest notes
- **Weaker than v1 in Togo, Chad, Samoa** (deltas above) β€” the headline
limitation of this release. Armenia, Georgia and Myanmar measure slightly
below v1 (statistically par).
- **Women with non-Latin-script names remain weak**: on Persian-script and
Cyrillic names female accuracy stays around 29–35 %, as low as 19.9 % on
Persian-script names in the fresh sets. The aggregate gains do not fix this.
- **Arabic and Hebrew: no gain.** Transliteration does not help these scripts
(measured: routing leaves them at the native branch); v2 is at best par with
v1 there. Thai likewise stays native.
- **The model does not condition gender on country**: e.g. "Andrea Costa" is
classified F with p(M) = 0.44 (Andrea is a male name in Italian, female in
Spanish).
- **No built-in abstention**: the model always answers. Probabilities are
temperature-calibrated per network; the 0.60 threshold was fixed before the
fresh-set verification. For analyses, use `p_male` (and a margin you choose)
rather than the hard label.
- **Binary output only**: the model returns M/F. It cannot represent
non-binary identities, and name-inferred gender is a *perceived-attribute
proxy*, not a statement about any person's identity.
- Bench numbers are specific to the named frozen benches; several are
Wikidata-derived and may share source characteristics with the training
corpora (the independent 226-v2 bench was built with explicit exclusions;
the old 226 bench overlaps v1 training by 34 %, as declared above).
## Data provenance
- Trained on **aggregated public name records with gender and country signals**
(Wikidata-derived corpora, Olympian records, per-country name lists),
internally reviewed. Training corpora, splits and recipes are **private and
not redistributed** with this package.
- **No LLM-generated labels were used in the training of this model.** An
LLM second-opinion cascade was explored in research and measured, but it is
**not part of this release**.
- A small fraction of the network-B corpus (**8,708 rows, β‰ˆ 1.21 %**) comes
from a GFDL-1.3-licensed public dataset; this is disclosed for transparency
about obligations.
- WGND and IPUMS were never used in training.
- Bench sources: Wikidata (CC0), Olympian and public record compilations;
bench row-level data is not redistributed.
## Licence
- **Weights and model artifacts** (everything under `genderize_fuso_v2/`):
**CC BY-NC 4.0** β€” see `LICENSE`; commercial use requires a separate
licence (see `COMMERCIAL_USE.md`).
- **Inference code** (`inference.py`, `inference_fuso.py`): **MIT** β€” see
`LICENSE-CODE`.
- Runtime dependency `anyascii`: ISC licence (third-party).
## Intended use
Aggregate, statistical analyses: bibliometrics and authorship studies,
name-collection demography, dataset documentation. **Not** intended for
decisions about individual people (hiring, credit, identity verification,
content moderation), nor as ground truth about anyone's gender.
## Version history
- **v2** (this release): fused two-network model, script-aware routing,
country prior, per-country selector. 14 above, 94 par, 1 below at β‰₯ 400 rows
(3 below at β‰₯ 200) on the independent 226-v2 bench (see above).
- **v1** (`cora` / `ultra`): preserved in this repository at git tag **`v1`**.
---
Support this work: https://github.com/sponsors/sheppard94g