kyaraembed / README.md
nonmetal's picture
KyaraEmbed v1.0
25bb09b
|
Raw History Blame Contribute Delete
2.72 kB
---
license: cc-by-nc-4.0
base_model: litagin/anime_speaker_embedding_ecapa_tdnn_groupnorm
pipeline_tag: audio-classification
tags:
- speaker-verification
- character-identity
- cross-lingual
- anime
- speech
language:
- ja
- en
---
# KyaraEmbed
An encoder for anime **character** identity rather than speaker identity. Given an audio clip, it
returns a 192-dimensional embedding whose cosine similarity follows the character across a change
of performer, and separates characters who share one.
Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was
built for live in the [KyaraBench repository](https://github.com/sizigi/kyarabench).
## Checkpoint
| File | Size | Notes |
|---|---|---|
| `kyaraembed_v1.0.pt` | 83 MB | ECAPA-TDNN, 20.8M parameters, 192-d output |
| `kyaraembed_v1.0.pt.sha256` | — | Checksum |
| `LICENSE-MODEL`, `NOTICE` | — | Licence and the upstream attribution chain |
## How to use
```bash
pip install torch soundfile librosa anime-speaker-embedding
```
```python
import numpy as np, torch, soundfile as sf, librosa
from anime_speaker_embedding import AnimeSpeakerEmbedding
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AnimeSpeakerEmbedding(device=device, variant="char")
model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device))
model.eval()
def load(path, seconds=10.0, sr=16000):
"""Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded."""
wav, file_sr = sf.read(path, dtype="float32", always_2d=True)
wav = wav.mean(axis=1)
if file_sr != sr:
wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr)
out = np.zeros(int(seconds * sr), dtype="float32")
wav = wav[:len(out)]
out[:len(wav)] = wav
return out
x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device)
with torch.no_grad():
z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
print(f"similarity: {float(z[0] @ z[1]):.3f}")
```
Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to
be in its batch changes its embedding; padding to a fixed length does not.
For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine:
```python
def prototype(paths):
x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device)
with torch.no_grad():
z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
return torch.nn.functional.normalize(z.mean(0), dim=-1)
```
## Licence
CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting
non-commercial research only. See `LICENSE-MODEL` and `NOTICE`.