KyaraEmbed

An encoder for anime character identity rather than speaker identity. Given an audio clip, it returns a 192-dimensional embedding whose cosine similarity follows the character across a change of performer, and separates characters who share one.

Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was built for live in the KyaraBench repository.

Checkpoint

File Size Notes
kyaraembed_v1.0.pt 83 MB ECAPA-TDNN, 20.8M parameters, 192-d output
kyaraembed_v1.0.pt.sha256 โ€” Checksum
LICENSE-MODEL, NOTICE โ€” Licence and the upstream attribution chain

How to use

pip install torch soundfile librosa anime-speaker-embedding
import numpy as np, torch, soundfile as sf, librosa
from anime_speaker_embedding import AnimeSpeakerEmbedding

device = "cuda" if torch.cuda.is_available() else "cpu"
model = AnimeSpeakerEmbedding(device=device, variant="char")
model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device))
model.eval()

def load(path, seconds=10.0, sr=16000):
    """Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded."""
    wav, file_sr = sf.read(path, dtype="float32", always_2d=True)
    wav = wav.mean(axis=1)
    if file_sr != sr:
        wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr)
    out = np.zeros(int(seconds * sr), dtype="float32")
    wav = wav[:len(out)]
    out[:len(wav)] = wav
    return out

x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device)
with torch.no_grad():
    z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
print(f"similarity: {float(z[0] @ z[1]):.3f}")

Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to be in its batch changes its embedding; padding to a fixed length does not.

For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine:

def prototype(paths):
    x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device)
    with torch.no_grad():
        z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
    return torch.nn.functional.normalize(z.mean(0), dim=-1)

Licence

CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting non-commercial research only. See LICENSE-MODEL and NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for spellbrush/kyaraembed