campplus-gender-classifier

Estimates whether a voice sounds female or male from a few turns of 16 kHz speech, with a calibrated confidence meant for gating decisions. It predicts perceived voice characteristics, not a person's gender identity.

  • Base: WeSpeaker CAM++ (VoxCeleb, Apache-2.0); fine-tuned: xvector.block3, xvector.transit3, xvector.out_nonlinear, xvector.dense, head
  • Output: p = sigmoid(mean turn logit / T), T = 0.820; confidence = |2p - 1|
  • Training audio augmented with noise, babble, an opposite-sex interferer, gain/clipping and an 8 kHz mu-law phone path
  • Sources: edacc, kathbath_hi, kathbath_ml, libritts, vctk (speaker-disjoint fit / calibration / test)

Gate on simulated calls (held-out speakers)

confidence >= answers on wrong
0.8 90.8% 2.40%
0.9 86.1% 1.97%
0.95 81.0% 1.40%
0.98 71.7% 1.17%

The test labels include known LibriTTS speaker-metadata errors (speakers whose recorded gender is flipped, confirmed by listening), so these error rates are upper bounds. Per-source and per-condition results: metrics.json.

Use

from inference import GenderClassifierModel
model = GenderClassifierModel(".")
model.predict([turn1_int16, turn2_int16])   # -> (p_female, confidence) or None

PyTorch weights for further fine-tuning: model.safetensors, loaded with model.GenderClassifier.from_folder(".").

Training data

LibriTTS-R (gender from parler-tts speaker descriptions), VCTK, EdAcc and Kathbath (Hindi, Malayalam). Each dataset is under its own licence; see its page.

Downloads last month
12
Safetensors
Model size
7.26M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AswanthCManoj/campplus-gender-classifier

Quantized
(1)
this model