|
Download README.md from spellbrush/kyaraembed: direct link, hf CLI and curl.
- Browser
- Download file 2.72 kB
-
https://huggingface.co/spellbrush/kyaraembed/resolve/main/README.md
- Command line
-
hf download hf://spellbrush/kyaraembed/README.md
-
curl -L -o README.md https://huggingface.co/spellbrush/kyaraembed/resolve/main/README.md
2.72 kB
| license: cc-by-nc-4.0 | |
| base_model: litagin/anime_speaker_embedding_ecapa_tdnn_groupnorm | |
| pipeline_tag: audio-classification | |
| tags: | |
| - speaker-verification | |
| - character-identity | |
| - cross-lingual | |
| - anime | |
| - speech | |
| language: | |
| - ja | |
| - en | |
| # KyaraEmbed | |
| An encoder for anime **character** identity rather than speaker identity. Given an audio clip, it | |
| returns a 192-dimensional embedding whose cosine similarity follows the character across a change | |
| of performer, and separates characters who share one. | |
| Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was | |
| built for live in the [KyaraBench repository](https://github.com/sizigi/kyarabench). | |
| ## Checkpoint | |
| | File | Size | Notes | | |
| |---|---|---| | |
| | `kyaraembed_v1.0.pt` | 83 MB | ECAPA-TDNN, 20.8M parameters, 192-d output | | |
| | `kyaraembed_v1.0.pt.sha256` | — | Checksum | | |
| | `LICENSE-MODEL`, `NOTICE` | — | Licence and the upstream attribution chain | | |
| ## How to use | |
| ```bash | |
| pip install torch soundfile librosa anime-speaker-embedding | |
| ``` | |
| ```python | |
| import numpy as np, torch, soundfile as sf, librosa | |
| from anime_speaker_embedding import AnimeSpeakerEmbedding | |
| device = "cuda" if torch.cuda.is_available() else "cpu" | |
| model = AnimeSpeakerEmbedding(device=device, variant="char") | |
| model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device)) | |
| model.eval() | |
| def load(path, seconds=10.0, sr=16000): | |
| """Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded.""" | |
| wav, file_sr = sf.read(path, dtype="float32", always_2d=True) | |
| wav = wav.mean(axis=1) | |
| if file_sr != sr: | |
| wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr) | |
| out = np.zeros(int(seconds * sr), dtype="float32") | |
| wav = wav[:len(out)] | |
| out[:len(wav)] = wav | |
| return out | |
| x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device) | |
| with torch.no_grad(): | |
| z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1) | |
| print(f"similarity: {float(z[0] @ z[1]):.3f}") | |
| ``` | |
| Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to | |
| be in its batch changes its embedding; padding to a fixed length does not. | |
| For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine: | |
| ```python | |
| def prototype(paths): | |
| x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device) | |
| with torch.no_grad(): | |
| z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1) | |
| return torch.nn.functional.normalize(z.mean(0), dim=-1) | |
| ``` | |
| ## Licence | |
| CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting | |
| non-commercial research only. See `LICENSE-MODEL` and `NOTICE`. | |