--- license: cc-by-nc-4.0 base_model: litagin/anime_speaker_embedding_ecapa_tdnn_groupnorm pipeline_tag: audio-classification tags: - speaker-verification - character-identity - cross-lingual - anime - speech language: - ja - en --- # KyaraEmbed An encoder for anime **character** identity rather than speaker identity. Given an audio clip, it returns a 192-dimensional embedding whose cosine similarity follows the character across a change of performer, and separates characters who share one. Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was built for live in the [KyaraBench repository](https://github.com/sizigi/kyarabench). ## Checkpoint | File | Size | Notes | |---|---|---| | `kyaraembed_v1.0.pt` | 83 MB | ECAPA-TDNN, 20.8M parameters, 192-d output | | `kyaraembed_v1.0.pt.sha256` | — | Checksum | | `LICENSE-MODEL`, `NOTICE` | — | Licence and the upstream attribution chain | ## How to use ```bash pip install torch soundfile librosa anime-speaker-embedding ``` ```python import numpy as np, torch, soundfile as sf, librosa from anime_speaker_embedding import AnimeSpeakerEmbedding device = "cuda" if torch.cuda.is_available() else "cpu" model = AnimeSpeakerEmbedding(device=device, variant="char") model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device)) model.eval() def load(path, seconds=10.0, sr=16000): """Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded.""" wav, file_sr = sf.read(path, dtype="float32", always_2d=True) wav = wav.mean(axis=1) if file_sr != sr: wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr) out = np.zeros(int(seconds * sr), dtype="float32") wav = wav[:len(out)] out[:len(wav)] = wav return out x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device) with torch.no_grad(): z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1) print(f"similarity: {float(z[0] @ z[1]):.3f}") ``` Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to be in its batch changes its embedding; padding to a fixed length does not. For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine: ```python def prototype(paths): x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device) with torch.no_grad(): z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1) return torch.nn.functional.normalize(z.mean(0), dim=-1) ``` ## Licence CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting non-commercial research only. See `LICENSE-MODEL` and `NOTICE`.