Genderize v2
Genderize v2 estimates a binary gender (M/F) and a likely country (226
ISO-3166 alpha-2 codes) from a personal name, with probabilities. It is a
character-level classifier designed for CPU inference β no tokenizer, no
vocabulary file. It supersedes genderize v1 (cora/ultra), which remains
available in this repository under the git tag v1.
Read the limitations before using this model. It is measurably weaker than v1 in Togo, Chad and Samoa, and it does not improve Arabic and Hebrew. It is intended for aggregate statistical analyses, not for decisions about individual people.
Package contents
| File | What it is | Licence |
|---|---|---|
genderize_fuso_v2/fuso.json |
fusion configuration (thresholds, weights) | CC BY-NC 4.0 |
genderize_fuso_v2/v1/ |
internal network A: the v1 ultra weights, configs, calibration, class map |
CC BY-NC 4.0 |
genderize_fuso_v2/v2/ |
internal network B (codepoint-level, distilled lineage), weights, configs, calibration, class map | CC BY-NC 4.0 |
genderize_fuso_v2/country_prior.json |
empirical country prior used by the country head | CC BY-NC 4.0 |
genderize_fuso_v2/dom113_rimedio.json |
configuration of the Latin-script country-trait rule (see below) | CC BY-NC 4.0 |
genderize_fuso_v2/SHA256SUMS |
SHA-256 of every file in the model directory | CC BY-NC 4.0 |
inference.py |
inference module for network A (byte-level model) | MIT |
inference_fuso.py |
the Genderize v2 class GenderizeFuso (imports inference.py) |
MIT |
examples.txt |
example outputs reproduced with exactly these files | β |
LICENSE |
CC BY-NC 4.0 (weights and model artifacts) | β |
LICENSE-CODE |
MIT (inference code) | β |
COMMERCIAL_USE.md |
what counts as non-commercial use, and how to obtain a commercial licence | β |
Training data is not distributed with this package (see Data provenance).
Quickstart
Requires Python 3.10+, torch, numpy, and anyascii (ISC licence; only
invoked for non-Latin scripts that benefit from transliteration).
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("textpie/genderize")
sys.path.insert(0, path) # inference_fuso.py + inference.py live at repo root
from inference_fuso import GenderizeFuso # the class documented in the file header
model = GenderizeFuso(f"{path}/genderize_fuso_v2")
print(model.predict(["Andrea Rossi", "Yuki Tanaka"]))
# -> [{'gender': 'M', 'p_male': 0.7937, 'country': 'IT', 'p_country': 0.8262, 'top5': ['IT', 'FR', 'US', 'CH', 'DE']},
# {'gender': 'F', 'p_male': 0.5037, 'country': 'JP', 'p_country': 0.9809, 'top5': ['JP', 'US', 'ID', 'DE', 'CN']}]
predict() returns one dict per name: gender ('M' if p(M) β₯ 0.60, else
'F'), p_male, country (argmax of the adjusted country distribution),
p_country, and top5 (top-5 country codes).
How it works
- Two internal networks, one fused output. Network A is the v1
ultrabyte-level CNN (UTF-8 bytes, 48-byte truncation). Network B reads the name as Unicode codepoints (32 max, its own 226-country map). Gender probability is the plain average of the two calibrated p(M); the country distribution is 0.7 Γ A + 0.3 Γ B. - Calibration and threshold. Each network ships per-head temperature calibration files. The operating threshold M β₯ 0.60 was chosen on three frozen benches and then verified β without retuning β on six fresh, never-touched sets.
- Script-aware routing (T1). The Unicode script of the name is detected
from character blocks; for measured scripts the model classifies the native
form, an
anyasciitransliteration, or the average of both (rule fixed on a validation half only, no learned parameters). Latin and unmapped scripts stay native. The country head always reads the original form. - Country prior. The country distribution is adjusted with a logit term
log p(c) β 0.75 Β· log Ο(c), whereΟis the empirical prior incountry_prior.json(coefficient fixed on validation only). - Per-country gender selector + Latin-script trait rule. For a small set
of countries where v2.2 measured regressions (AM, GE, MM, TG, TD, WS), part
of the gender decision falls back to the v1 branch under rules fixed on old
benches and on a frozen validation bench β never on the 226-v2 test bench
(configuration in
dom113_rimedio.json).
Results (gender accuracy, measured)
v1 = genderize-ultra (calibrated, threshold 0.50). v2 = this release
(threshold 0.60). All numbers are point measurements on frozen benches; they
do not generalize beyond the named benches.
Independent bench: 226-countries v2 (50,258 rows, Latin script)
Built with explicit exclusions against the training corpora and all earlier benches; no overlap with training (full = residual by construction). Paired bootstrap, 2,000 resamples, seed 20260927.
| Group | n | v2 | v1 |
|---|---|---|---|
| All | 50,258 | 94.08 % | 93.71 % |
| Women (F) | 20,648 | 92.98 % | 90.30 % |
| Men (M) | 29,610 | 94.86 % | 96.09 % |
Per country (countries with β₯ 400 bench rows): 14 countries significantly above v1, 94 statistically par, 1 below. With the pre-registered criterion (β₯ 200 rows): 14 above, 106 par, 3 below. Countries above v1: AZ, DZ, EG, ET, HK, JM, LK, MY, NG, NP, SG, SY, UY, ZA. "Par" means the confidence interval includes zero β it is not proof of equivalence; no multiple-comparison correction is applied.
Where v2 is worse than v1 (all measured deltas, IC95 in percentage points)
| Country | n | Ξ (v2 β v1) | IC95 |
|---|---|---|---|
| Togo (TG) | 400 | β2.00 | [β3.50, β0.50] |
| Chad (TD) | 207 | β2.42 | [β4.83, β0.48] |
| Samoa (WS) | 210 | β4.76 | [β9.05, β0.95] |
| Armenia (AM) | 400 | β0.25 | [β0.75, 0.00] |
| Georgia (GE) | 400 | β0.25 | [β0.75, 0.00] |
| Myanmar (MM) | 400 | β1.25 | [β3.50, +0.75] |
TG is the only country below v1 at the β₯ 400-row quota; TD and WS are also below at the pre-registered β₯ 200-row criterion; AM, GE and MM measure statistically par. Without the Latin-script trait rule, deltas down to β7.3 pp were measured on the same bench. If your analysis focuses on these countries, use v1.
Other frozen benches
| Bench | n | v2 | v1 | Notes |
|---|---|---|---|---|
| Native scripts (with T1 routing) | 16,945 | 88.70 % | 83.84 % | release path (v2.2 verdict) |
| Native scripts + routing (T1) | 6,072 | 85.82 % | 74.31 % | half-test split; largest gains: Georgian 50.9β92.3, Bengali 62.5β87.7, Armenian 67.5β90.8, Kazakh 93.1β97.5, Hindi 83.8β87.3 |
| East-Asian | 5,199 | 80.96 % | 79.75 % | women 79.3 vs 69.3 |
| Old 226 bench (Wikidata) | 39,164 | 92.63 % | 92.14 % | not independent: 34 % of this bench is inside v1's training data (v1 residual: 90.41 %); kept for continuity with v1's history |
Six fresh, never-touched sets (v2 / v1): Olympians 96.20 / 96.07; Vietnamese 80.32 / 76.58; Tajik-Latin 79.09 / 76.13; Tajik-Cyrillic 72.12 / 67.65; Persian-Latin 77.57 / 74.07; Persian-script 67.29 / 64.59.
Country (nationality) output
The country head is a coarse signal: on the old 226 bench, top-1 β 28.4 %
(v1: 28.1 %; v1 top-5 β 56.3 %). Use top5 and the probability, not the
argmax alone.
Persian dedicated branch: tested and rejected
A dedicated Persian transliteration branch (beam over k readings, < 5 MB) was integrated and measured on the only available native-Persian bench (n = 1,000; not fresh β already used by earlier experiments): +2.5 pp over the direct branch, IC95 [β0.4, +5.6] β the interval crosses zero, so the gain is not confirmed and the branch is not part of this release. It will be reconsidered only if an independent, fresh native-Persian bench becomes available (none found as of 2026-09-28).
Speed
198.8 names/s with script routing (189.6 without), single CPU, 8 threads, batch 256, measured back-to-back on the bench machine. Throughput varies with hardware and load; no latency promise is made.
Limitations and honest notes
- Weaker than v1 in Togo, Chad, Samoa (deltas above) β the headline limitation of this release. Armenia, Georgia and Myanmar measure slightly below v1 (statistically par).
- Women with non-Latin-script names remain weak: on Persian-script and Cyrillic names female accuracy stays around 29β35 %, as low as 19.9 % on Persian-script names in the fresh sets. The aggregate gains do not fix this.
- Arabic and Hebrew: no gain. Transliteration does not help these scripts (measured: routing leaves them at the native branch); v2 is at best par with v1 there. Thai likewise stays native.
- The model does not condition gender on country: e.g. "Andrea Costa" is classified F with p(M) = 0.44 (Andrea is a male name in Italian, female in Spanish).
- No built-in abstention: the model always answers. Probabilities are
temperature-calibrated per network; the 0.60 threshold was fixed before the
fresh-set verification. For analyses, use
p_male(and a margin you choose) rather than the hard label. - Binary output only: the model returns M/F. It cannot represent non-binary identities, and name-inferred gender is a perceived-attribute proxy, not a statement about any person's identity.
- Bench numbers are specific to the named frozen benches; several are Wikidata-derived and may share source characteristics with the training corpora (the independent 226-v2 bench was built with explicit exclusions; the old 226 bench overlaps v1 training by 34 %, as declared above).
Data provenance
- Trained on aggregated public name records with gender and country signals (Wikidata-derived corpora, Olympian records, per-country name lists), internally reviewed. Training corpora, splits and recipes are private and not redistributed with this package.
- No LLM-generated labels were used in the training of this model. An LLM second-opinion cascade was explored in research and measured, but it is not part of this release.
- A small fraction of the network-B corpus (8,708 rows, β 1.21 %) comes from a GFDL-1.3-licensed public dataset; this is disclosed for transparency about obligations.
- WGND and IPUMS were never used in training.
- Bench sources: Wikidata (CC0), Olympian and public record compilations; bench row-level data is not redistributed.
Licence
- Weights and model artifacts (everything under
genderize_fuso_v2/): CC BY-NC 4.0 β seeLICENSE; commercial use requires a separate licence (seeCOMMERCIAL_USE.md). - Inference code (
inference.py,inference_fuso.py): MIT β seeLICENSE-CODE. - Runtime dependency
anyascii: ISC licence (third-party).
Intended use
Aggregate, statistical analyses: bibliometrics and authorship studies, name-collection demography, dataset documentation. Not intended for decisions about individual people (hiring, credit, identity verification, content moderation), nor as ground truth about anyone's gender.
Version history
- v2 (this release): fused two-network model, script-aware routing, country prior, per-country selector. 14 above, 94 par, 1 below at β₯ 400 rows (3 below at β₯ 200) on the independent 226-v2 bench (see above).
- v1 (
cora/ultra): preserved in this repository at git tagv1.
Support this work: https://github.com/sponsors/sheppard94g