|
Download README.md from textpie/genderize: direct link, hf CLI and curl.
- Browser
- Download file 11.8 kB
-
https://huggingface.co/textpie/genderize/resolve/main/README.md
- Command line
-
hf download hf://textpie/genderize/README.md
-
curl -L -o README.md https://huggingface.co/textpie/genderize/resolve/main/README.md
11.8 kB
| language: | |
| - multilingual | |
| license: cc-by-nc-4.0 | |
| library_name: pytorch | |
| tags: | |
| - text-classification | |
| - gender-detection | |
| - nationality | |
| - name-analysis | |
| - character-level | |
| - cpu | |
| pipeline_tag: text-classification | |
| # Genderize v2 | |
| Genderize v2 estimates a **binary gender** (M/F) and a **likely country** (226 | |
| ISO-3166 alpha-2 codes) from a personal name, with probabilities. It is a | |
| character-level classifier designed for **CPU inference** β no tokenizer, no | |
| vocabulary file. It supersedes `genderize v1 (cora/ultra)`, which remains | |
| available in this repository under the git tag **`v1`**. | |
| > **Read the limitations before using this model.** It is measurably **weaker | |
| > than v1 in Togo, Chad and Samoa**, and it does not improve Arabic and Hebrew. | |
| > It is intended for **aggregate statistical analyses**, not for decisions | |
| > about individual people. | |
| ## Package contents | |
| | File | What it is | Licence | | |
| |---|---|---| | |
| | `genderize_fuso_v2/fuso.json` | fusion configuration (thresholds, weights) | CC BY-NC 4.0 | | |
| | `genderize_fuso_v2/v1/` | internal network A: the v1 `ultra` weights, configs, calibration, class map | CC BY-NC 4.0 | | |
| | `genderize_fuso_v2/v2/` | internal network B (codepoint-level, distilled lineage), weights, configs, calibration, class map | CC BY-NC 4.0 | | |
| | `genderize_fuso_v2/country_prior.json` | empirical country prior used by the country head | CC BY-NC 4.0 | | |
| | `genderize_fuso_v2/dom113_rimedio.json` | configuration of the Latin-script country-trait rule (see below) | CC BY-NC 4.0 | | |
| | `genderize_fuso_v2/SHA256SUMS` | SHA-256 of every file in the model directory | CC BY-NC 4.0 | | |
| | `inference.py` | inference module for network A (byte-level model) | MIT | | |
| | `inference_fuso.py` | the Genderize v2 class `GenderizeFuso` (imports `inference.py`) | MIT | | |
| | `examples.txt` | example outputs reproduced with exactly these files | β | | |
| | `LICENSE` | CC BY-NC 4.0 (weights and model artifacts) | β | | |
| | `LICENSE-CODE` | MIT (inference code) | β | | |
| | `COMMERCIAL_USE.md` | what counts as non-commercial use, and how to obtain a commercial licence | β | | |
| Training data is **not** distributed with this package (see *Data provenance*). | |
| ## Quickstart | |
| Requires Python 3.10+, `torch`, `numpy`, and `anyascii` (ISC licence; only | |
| invoked for non-Latin scripts that benefit from transliteration). | |
| ```python | |
| import sys | |
| from huggingface_hub import snapshot_download | |
| path = snapshot_download("textpie/genderize") | |
| sys.path.insert(0, path) # inference_fuso.py + inference.py live at repo root | |
| from inference_fuso import GenderizeFuso # the class documented in the file header | |
| model = GenderizeFuso(f"{path}/genderize_fuso_v2") | |
| print(model.predict(["Andrea Rossi", "Yuki Tanaka"])) | |
| # -> [{'gender': 'M', 'p_male': 0.7937, 'country': 'IT', 'p_country': 0.8262, 'top5': ['IT', 'FR', 'US', 'CH', 'DE']}, | |
| # {'gender': 'F', 'p_male': 0.5037, 'country': 'JP', 'p_country': 0.9809, 'top5': ['JP', 'US', 'ID', 'DE', 'CN']}] | |
| ``` | |
| `predict()` returns one dict per name: `gender` (`'M'` if p(M) β₯ 0.60, else | |
| `'F'`), `p_male`, `country` (argmax of the adjusted country distribution), | |
| `p_country`, and `top5` (top-5 country codes). | |
| ## How it works | |
| - **Two internal networks, one fused output.** Network A is the v1 `ultra` | |
| byte-level CNN (UTF-8 bytes, 48-byte truncation). Network B reads the name | |
| as Unicode codepoints (32 max, its own 226-country map). Gender probability | |
| is the plain average of the two calibrated p(M); the country distribution is | |
| 0.7 Γ A + 0.3 Γ B. | |
| - **Calibration and threshold.** Each network ships per-head temperature | |
| calibration files. The operating threshold M β₯ **0.60** was chosen on three | |
| frozen benches and then verified β without retuning β on six fresh, | |
| never-touched sets. | |
| - **Script-aware routing (T1).** The Unicode script of the name is detected | |
| from character blocks; for measured scripts the model classifies the native | |
| form, an `anyascii` transliteration, or the average of both (rule fixed on a | |
| validation half only, no learned parameters). Latin and unmapped scripts | |
| stay native. The country head always reads the original form. | |
| - **Country prior.** The country distribution is adjusted with a logit term | |
| `log p(c) β 0.75 Β· log Ο(c)`, where `Ο` is the empirical prior in | |
| `country_prior.json` (coefficient fixed on validation only). | |
| - **Per-country gender selector + Latin-script trait rule.** For a small set | |
| of countries where v2.2 measured regressions (AM, GE, MM, TG, TD, WS), part | |
| of the gender decision falls back to the v1 branch under rules fixed on old | |
| benches and on a frozen validation bench β never on the 226-v2 test bench | |
| (configuration in `dom113_rimedio.json`). | |
| ## Results (gender accuracy, measured) | |
| v1 = `genderize-ultra` (calibrated, threshold 0.50). v2 = this release | |
| (threshold 0.60). All numbers are point measurements on frozen benches; they | |
| do not generalize beyond the named benches. | |
| ### Independent bench: 226-countries v2 (50,258 rows, Latin script) | |
| Built with explicit exclusions against the training corpora and all earlier | |
| benches; **no overlap with training** (full = residual by construction). | |
| Paired bootstrap, 2,000 resamples, seed 20260927. | |
| | Group | n | v2 | v1 | | |
| |---|---:|---:|---:| | |
| | All | 50,258 | **94.08 %** | 93.71 % | | |
| | Women (F) | 20,648 | **92.98 %** | 90.30 % | | |
| | Men (M) | 29,610 | 94.86 % | **96.09 %** | | |
| Per country (countries with β₯ 400 bench rows): **14 countries significantly | |
| above v1, 94 statistically par, 1 below**. With the pre-registered criterion | |
| (β₯ 200 rows): **14 above, 106 par, 3 below**. Countries above v1: AZ, DZ, | |
| EG, ET, HK, JM, LK, MY, NG, NP, SG, SY, UY, ZA. "Par" means the confidence | |
| interval includes zero β it is not proof of equivalence; no multiple-comparison | |
| correction is applied. | |
| ### Where v2 is **worse than v1** (all measured deltas, IC95 in percentage points) | |
| | Country | n | Ξ (v2 β v1) | IC95 | | |
| |---|---:|---:|---| | |
| | **Togo (TG)** | 400 | **β2.00** | [β3.50, β0.50] | | |
| | **Chad (TD)** | 207 | **β2.42** | [β4.83, β0.48] | | |
| | **Samoa (WS)** | 210 | **β4.76** | [β9.05, β0.95] | | |
| | Armenia (AM) | 400 | β0.25 | [β0.75, 0.00] | | |
| | Georgia (GE) | 400 | β0.25 | [β0.75, 0.00] | | |
| | Myanmar (MM) | 400 | β1.25 | [β3.50, +0.75] | | |
| TG is the only country below v1 at the β₯ 400-row quota; TD and WS are also | |
| below at the pre-registered β₯ 200-row criterion; AM, GE and MM measure | |
| statistically par. Without the Latin-script trait rule, deltas down to | |
| β7.3 pp were measured on the same bench. **If your | |
| analysis focuses on these countries, use v1.** | |
| ### Other frozen benches | |
| | Bench | n | v2 | v1 | Notes | | |
| |---|---:|---:|---:|---| | |
| | Native scripts (with T1 routing) | 16,945 | **88.70 %** | 83.84 % | release path (v2.2 verdict) | | |
| | Native scripts + routing (T1) | 6,072 | **85.82 %** | 74.31 % | half-test split; largest gains: Georgian 50.9β92.3, Bengali 62.5β87.7, Armenian 67.5β90.8, Kazakh 93.1β97.5, Hindi 83.8β87.3 | | |
| | East-Asian | 5,199 | **80.96 %** | 79.75 % | women 79.3 vs 69.3 | | |
| | Old 226 bench (Wikidata) | 39,164 | 92.63 % | 92.14 % | **not independent: 34 % of this bench is inside v1's training data** (v1 residual: 90.41 %); kept for continuity with v1's history | | |
| Six fresh, never-touched sets (v2 / v1): Olympians 96.20 / 96.07; Vietnamese | |
| 80.32 / 76.58; Tajik-Latin 79.09 / 76.13; Tajik-Cyrillic 72.12 / 67.65; | |
| Persian-Latin 77.57 / 74.07; Persian-script 67.29 / 64.59. | |
| ### Country (nationality) output | |
| The country head is a **coarse signal**: on the old 226 bench, top-1 β 28.4 % | |
| (v1: 28.1 %; v1 top-5 β 56.3 %). Use `top5` and the probability, not the | |
| argmax alone. | |
| ### Persian dedicated branch: tested and **rejected** | |
| A dedicated Persian transliteration branch (beam over k readings, < 5 MB) was | |
| integrated and measured on the only available native-Persian bench (n = 1,000; | |
| **not fresh** β already used by earlier experiments): +2.5 pp over the direct | |
| branch, IC95 [β0.4, +5.6] β the interval crosses zero, so the gain is **not | |
| confirmed** and the branch is **not part of this release**. It will be | |
| reconsidered only if an independent, fresh native-Persian bench becomes | |
| available (none found as of 2026-09-28). | |
| ### Speed | |
| 198.8 names/s with script routing (189.6 without), single CPU, 8 threads, | |
| batch 256, measured back-to-back on the bench machine. Throughput varies with | |
| hardware and load; no latency promise is made. | |
| ## Limitations and honest notes | |
| - **Weaker than v1 in Togo, Chad, Samoa** (deltas above) β the headline | |
| limitation of this release. Armenia, Georgia and Myanmar measure slightly | |
| below v1 (statistically par). | |
| - **Women with non-Latin-script names remain weak**: on Persian-script and | |
| Cyrillic names female accuracy stays around 29β35 %, as low as 19.9 % on | |
| Persian-script names in the fresh sets. The aggregate gains do not fix this. | |
| - **Arabic and Hebrew: no gain.** Transliteration does not help these scripts | |
| (measured: routing leaves them at the native branch); v2 is at best par with | |
| v1 there. Thai likewise stays native. | |
| - **The model does not condition gender on country**: e.g. "Andrea Costa" is | |
| classified F with p(M) = 0.44 (Andrea is a male name in Italian, female in | |
| Spanish). | |
| - **No built-in abstention**: the model always answers. Probabilities are | |
| temperature-calibrated per network; the 0.60 threshold was fixed before the | |
| fresh-set verification. For analyses, use `p_male` (and a margin you choose) | |
| rather than the hard label. | |
| - **Binary output only**: the model returns M/F. It cannot represent | |
| non-binary identities, and name-inferred gender is a *perceived-attribute | |
| proxy*, not a statement about any person's identity. | |
| - Bench numbers are specific to the named frozen benches; several are | |
| Wikidata-derived and may share source characteristics with the training | |
| corpora (the independent 226-v2 bench was built with explicit exclusions; | |
| the old 226 bench overlaps v1 training by 34 %, as declared above). | |
| ## Data provenance | |
| - Trained on **aggregated public name records with gender and country signals** | |
| (Wikidata-derived corpora, Olympian records, per-country name lists), | |
| internally reviewed. Training corpora, splits and recipes are **private and | |
| not redistributed** with this package. | |
| - **No LLM-generated labels were used in the training of this model.** An | |
| LLM second-opinion cascade was explored in research and measured, but it is | |
| **not part of this release**. | |
| - A small fraction of the network-B corpus (**8,708 rows, β 1.21 %**) comes | |
| from a GFDL-1.3-licensed public dataset; this is disclosed for transparency | |
| about obligations. | |
| - WGND and IPUMS were never used in training. | |
| - Bench sources: Wikidata (CC0), Olympian and public record compilations; | |
| bench row-level data is not redistributed. | |
| ## Licence | |
| - **Weights and model artifacts** (everything under `genderize_fuso_v2/`): | |
| **CC BY-NC 4.0** β see `LICENSE`; commercial use requires a separate | |
| licence (see `COMMERCIAL_USE.md`). | |
| - **Inference code** (`inference.py`, `inference_fuso.py`): **MIT** β see | |
| `LICENSE-CODE`. | |
| - Runtime dependency `anyascii`: ISC licence (third-party). | |
| ## Intended use | |
| Aggregate, statistical analyses: bibliometrics and authorship studies, | |
| name-collection demography, dataset documentation. **Not** intended for | |
| decisions about individual people (hiring, credit, identity verification, | |
| content moderation), nor as ground truth about anyone's gender. | |
| ## Version history | |
| - **v2** (this release): fused two-network model, script-aware routing, | |
| country prior, per-country selector. 14 above, 94 par, 1 below at β₯ 400 rows | |
| (3 below at β₯ 200) on the independent 226-v2 bench (see above). | |
| - **v1** (`cora` / `ultra`): preserved in this repository at git tag **`v1`**. | |
| --- | |
| Support this work: https://github.com/sponsors/sheppard94g | |