campplus-gender-classifier
Estimates whether a voice sounds female or male from a few turns of 16 kHz speech, with a calibrated confidence meant for gating decisions. It predicts perceived voice characteristics, not a person's gender identity.
- Base: WeSpeaker CAM++ (VoxCeleb, Apache-2.0); fine-tuned: xvector.block3, xvector.transit3, xvector.out_nonlinear, xvector.dense, head
- Output: p = sigmoid(mean turn logit / T), T = 0.820; confidence = |2p - 1|
- Training audio augmented with noise, babble, an opposite-sex interferer, gain/clipping and an 8 kHz mu-law phone path
- Sources: edacc, kathbath_hi, kathbath_ml, libritts, vctk (speaker-disjoint fit / calibration / test)
Gate on simulated calls (held-out speakers)
| confidence >= | answers on | wrong |
|---|---|---|
| 0.8 | 90.8% | 2.40% |
| 0.9 | 86.1% | 1.97% |
| 0.95 | 81.0% | 1.40% |
| 0.98 | 71.7% | 1.17% |
The test labels include known LibriTTS speaker-metadata errors (speakers whose recorded gender is flipped, confirmed
by listening), so these error rates are upper bounds. Per-source and per-condition results: metrics.json.
Use
from inference import GenderClassifierModel
model = GenderClassifierModel(".")
model.predict([turn1_int16, turn2_int16]) # -> (p_female, confidence) or None
PyTorch weights for further fine-tuning: model.safetensors, loaded with model.GenderClassifier.from_folder(".").
Training data
LibriTTS-R (gender from parler-tts speaker descriptions), VCTK, EdAcc and Kathbath (Hindi, Malayalam). Each dataset is under its own licence; see its page.
- Downloads last month
- 12
Model tree for AswanthCManoj/campplus-gender-classifier
Base model
Wespeaker/wespeaker-voxceleb-campplus