Genderize v2

Genderize v2 estimates a binary gender (M/F) and a likely country (226 ISO-3166 alpha-2 codes) from a personal name, with probabilities. It is a character-level classifier designed for CPU inference β€” no tokenizer, no vocabulary file. It supersedes genderize v1 (cora/ultra), which remains available in this repository under the git tag v1.

Read the limitations before using this model. It is measurably weaker than v1 in Togo, Chad and Samoa, and it does not improve Arabic and Hebrew. It is intended for aggregate statistical analyses, not for decisions about individual people.

Package contents

File What it is Licence
genderize_fuso_v2/fuso.json fusion configuration (thresholds, weights) CC BY-NC 4.0
genderize_fuso_v2/v1/ internal network A: the v1 ultra weights, configs, calibration, class map CC BY-NC 4.0
genderize_fuso_v2/v2/ internal network B (codepoint-level, distilled lineage), weights, configs, calibration, class map CC BY-NC 4.0
genderize_fuso_v2/country_prior.json empirical country prior used by the country head CC BY-NC 4.0
genderize_fuso_v2/dom113_rimedio.json configuration of the Latin-script country-trait rule (see below) CC BY-NC 4.0
genderize_fuso_v2/SHA256SUMS SHA-256 of every file in the model directory CC BY-NC 4.0
inference.py inference module for network A (byte-level model) MIT
inference_fuso.py the Genderize v2 class GenderizeFuso (imports inference.py) MIT
examples.txt example outputs reproduced with exactly these files β€”
LICENSE CC BY-NC 4.0 (weights and model artifacts) β€”
LICENSE-CODE MIT (inference code) β€”
COMMERCIAL_USE.md what counts as non-commercial use, and how to obtain a commercial licence β€”

Training data is not distributed with this package (see Data provenance).

Quickstart

Requires Python 3.10+, torch, numpy, and anyascii (ISC licence; only invoked for non-Latin scripts that benefit from transliteration).

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("textpie/genderize")
sys.path.insert(0, path)                     # inference_fuso.py + inference.py live at repo root

from inference_fuso import GenderizeFuso     # the class documented in the file header

model = GenderizeFuso(f"{path}/genderize_fuso_v2")
print(model.predict(["Andrea Rossi", "Yuki Tanaka"]))
# -> [{'gender': 'M', 'p_male': 0.7937, 'country': 'IT', 'p_country': 0.8262, 'top5': ['IT', 'FR', 'US', 'CH', 'DE']},
#     {'gender': 'F', 'p_male': 0.5037, 'country': 'JP', 'p_country': 0.9809, 'top5': ['JP', 'US', 'ID', 'DE', 'CN']}]

predict() returns one dict per name: gender ('M' if p(M) β‰₯ 0.60, else 'F'), p_male, country (argmax of the adjusted country distribution), p_country, and top5 (top-5 country codes).

How it works

  • Two internal networks, one fused output. Network A is the v1 ultra byte-level CNN (UTF-8 bytes, 48-byte truncation). Network B reads the name as Unicode codepoints (32 max, its own 226-country map). Gender probability is the plain average of the two calibrated p(M); the country distribution is 0.7 Γ— A + 0.3 Γ— B.
  • Calibration and threshold. Each network ships per-head temperature calibration files. The operating threshold M β‰₯ 0.60 was chosen on three frozen benches and then verified β€” without retuning β€” on six fresh, never-touched sets.
  • Script-aware routing (T1). The Unicode script of the name is detected from character blocks; for measured scripts the model classifies the native form, an anyascii transliteration, or the average of both (rule fixed on a validation half only, no learned parameters). Latin and unmapped scripts stay native. The country head always reads the original form.
  • Country prior. The country distribution is adjusted with a logit term log p(c) βˆ’ 0.75 Β· log Ο€(c), where Ο€ is the empirical prior in country_prior.json (coefficient fixed on validation only).
  • Per-country gender selector + Latin-script trait rule. For a small set of countries where v2.2 measured regressions (AM, GE, MM, TG, TD, WS), part of the gender decision falls back to the v1 branch under rules fixed on old benches and on a frozen validation bench β€” never on the 226-v2 test bench (configuration in dom113_rimedio.json).

Results (gender accuracy, measured)

v1 = genderize-ultra (calibrated, threshold 0.50). v2 = this release (threshold 0.60). All numbers are point measurements on frozen benches; they do not generalize beyond the named benches.

Independent bench: 226-countries v2 (50,258 rows, Latin script)

Built with explicit exclusions against the training corpora and all earlier benches; no overlap with training (full = residual by construction). Paired bootstrap, 2,000 resamples, seed 20260927.

Group n v2 v1
All 50,258 94.08 % 93.71 %
Women (F) 20,648 92.98 % 90.30 %
Men (M) 29,610 94.86 % 96.09 %

Per country (countries with β‰₯ 400 bench rows): 14 countries significantly above v1, 94 statistically par, 1 below. With the pre-registered criterion (β‰₯ 200 rows): 14 above, 106 par, 3 below. Countries above v1: AZ, DZ, EG, ET, HK, JM, LK, MY, NG, NP, SG, SY, UY, ZA. "Par" means the confidence interval includes zero β€” it is not proof of equivalence; no multiple-comparison correction is applied.

Where v2 is worse than v1 (all measured deltas, IC95 in percentage points)

Country n Ξ” (v2 βˆ’ v1) IC95
Togo (TG) 400 βˆ’2.00 [βˆ’3.50, βˆ’0.50]
Chad (TD) 207 βˆ’2.42 [βˆ’4.83, βˆ’0.48]
Samoa (WS) 210 βˆ’4.76 [βˆ’9.05, βˆ’0.95]
Armenia (AM) 400 βˆ’0.25 [βˆ’0.75, 0.00]
Georgia (GE) 400 βˆ’0.25 [βˆ’0.75, 0.00]
Myanmar (MM) 400 βˆ’1.25 [βˆ’3.50, +0.75]

TG is the only country below v1 at the β‰₯ 400-row quota; TD and WS are also below at the pre-registered β‰₯ 200-row criterion; AM, GE and MM measure statistically par. Without the Latin-script trait rule, deltas down to βˆ’7.3 pp were measured on the same bench. If your analysis focuses on these countries, use v1.

Other frozen benches

Bench n v2 v1 Notes
Native scripts (with T1 routing) 16,945 88.70 % 83.84 % release path (v2.2 verdict)
Native scripts + routing (T1) 6,072 85.82 % 74.31 % half-test split; largest gains: Georgian 50.9β†’92.3, Bengali 62.5β†’87.7, Armenian 67.5β†’90.8, Kazakh 93.1β†’97.5, Hindi 83.8β†’87.3
East-Asian 5,199 80.96 % 79.75 % women 79.3 vs 69.3
Old 226 bench (Wikidata) 39,164 92.63 % 92.14 % not independent: 34 % of this bench is inside v1's training data (v1 residual: 90.41 %); kept for continuity with v1's history

Six fresh, never-touched sets (v2 / v1): Olympians 96.20 / 96.07; Vietnamese 80.32 / 76.58; Tajik-Latin 79.09 / 76.13; Tajik-Cyrillic 72.12 / 67.65; Persian-Latin 77.57 / 74.07; Persian-script 67.29 / 64.59.

Country (nationality) output

The country head is a coarse signal: on the old 226 bench, top-1 β‰ˆ 28.4 % (v1: 28.1 %; v1 top-5 β‰ˆ 56.3 %). Use top5 and the probability, not the argmax alone.

Persian dedicated branch: tested and rejected

A dedicated Persian transliteration branch (beam over k readings, < 5 MB) was integrated and measured on the only available native-Persian bench (n = 1,000; not fresh β€” already used by earlier experiments): +2.5 pp over the direct branch, IC95 [βˆ’0.4, +5.6] β€” the interval crosses zero, so the gain is not confirmed and the branch is not part of this release. It will be reconsidered only if an independent, fresh native-Persian bench becomes available (none found as of 2026-09-28).

Speed

198.8 names/s with script routing (189.6 without), single CPU, 8 threads, batch 256, measured back-to-back on the bench machine. Throughput varies with hardware and load; no latency promise is made.

Limitations and honest notes

  • Weaker than v1 in Togo, Chad, Samoa (deltas above) β€” the headline limitation of this release. Armenia, Georgia and Myanmar measure slightly below v1 (statistically par).
  • Women with non-Latin-script names remain weak: on Persian-script and Cyrillic names female accuracy stays around 29–35 %, as low as 19.9 % on Persian-script names in the fresh sets. The aggregate gains do not fix this.
  • Arabic and Hebrew: no gain. Transliteration does not help these scripts (measured: routing leaves them at the native branch); v2 is at best par with v1 there. Thai likewise stays native.
  • The model does not condition gender on country: e.g. "Andrea Costa" is classified F with p(M) = 0.44 (Andrea is a male name in Italian, female in Spanish).
  • No built-in abstention: the model always answers. Probabilities are temperature-calibrated per network; the 0.60 threshold was fixed before the fresh-set verification. For analyses, use p_male (and a margin you choose) rather than the hard label.
  • Binary output only: the model returns M/F. It cannot represent non-binary identities, and name-inferred gender is a perceived-attribute proxy, not a statement about any person's identity.
  • Bench numbers are specific to the named frozen benches; several are Wikidata-derived and may share source characteristics with the training corpora (the independent 226-v2 bench was built with explicit exclusions; the old 226 bench overlaps v1 training by 34 %, as declared above).

Data provenance

  • Trained on aggregated public name records with gender and country signals (Wikidata-derived corpora, Olympian records, per-country name lists), internally reviewed. Training corpora, splits and recipes are private and not redistributed with this package.
  • No LLM-generated labels were used in the training of this model. An LLM second-opinion cascade was explored in research and measured, but it is not part of this release.
  • A small fraction of the network-B corpus (8,708 rows, β‰ˆ 1.21 %) comes from a GFDL-1.3-licensed public dataset; this is disclosed for transparency about obligations.
  • WGND and IPUMS were never used in training.
  • Bench sources: Wikidata (CC0), Olympian and public record compilations; bench row-level data is not redistributed.

Licence

  • Weights and model artifacts (everything under genderize_fuso_v2/): CC BY-NC 4.0 β€” see LICENSE; commercial use requires a separate licence (see COMMERCIAL_USE.md).
  • Inference code (inference.py, inference_fuso.py): MIT β€” see LICENSE-CODE.
  • Runtime dependency anyascii: ISC licence (third-party).

Intended use

Aggregate, statistical analyses: bibliometrics and authorship studies, name-collection demography, dataset documentation. Not intended for decisions about individual people (hiring, credit, identity verification, content moderation), nor as ground truth about anyone's gender.

Version history

  • v2 (this release): fused two-network model, script-aware routing, country prior, per-country selector. 14 above, 94 par, 1 below at β‰₯ 400 rows (3 below at β‰₯ 200) on the independent 226-v2 bench (see above).
  • v1 (cora / ultra): preserved in this repository at git tag v1.

Support this work: https://github.com/sponsors/sheppard94g

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support