KyaraEmbed
An encoder for anime character identity rather than speaker identity. Given an audio clip, it returns a 192-dimensional embedding whose cosine similarity follows the character across a change of performer, and separates characters who share one.
Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was built for live in the KyaraBench repository.
Checkpoint
| File | Size | Notes |
|---|---|---|
kyaraembed_v1.0.pt |
83 MB | ECAPA-TDNN, 20.8M parameters, 192-d output |
kyaraembed_v1.0.pt.sha256 |
โ | Checksum |
LICENSE-MODEL, NOTICE |
โ | Licence and the upstream attribution chain |
How to use
pip install torch soundfile librosa anime-speaker-embedding
import numpy as np, torch, soundfile as sf, librosa
from anime_speaker_embedding import AnimeSpeakerEmbedding
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AnimeSpeakerEmbedding(device=device, variant="char")
model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device))
model.eval()
def load(path, seconds=10.0, sr=16000):
"""Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded."""
wav, file_sr = sf.read(path, dtype="float32", always_2d=True)
wav = wav.mean(axis=1)
if file_sr != sr:
wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr)
out = np.zeros(int(seconds * sr), dtype="float32")
wav = wav[:len(out)]
out[:len(wav)] = wav
return out
x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device)
with torch.no_grad():
z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
print(f"similarity: {float(z[0] @ z[1]):.3f}")
Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to be in its batch changes its embedding; padding to a fixed length does not.
For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine:
def prototype(paths):
x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device)
with torch.no_grad():
z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
return torch.nn.functional.normalize(z.mean(0), dim=-1)
Licence
CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting
non-commercial research only. See LICENSE-MODEL and NOTICE.
Model tree for spellbrush/kyaraembed
Base model
speechbrain/spkrec-ecapa-voxceleb