kyaraembed / README.md
nonmetal's picture
KyaraEmbed v1.0
25bb09b
|
Raw History Blame Contribute Delete
2.72 kB
metadata
license: cc-by-nc-4.0
base_model: litagin/anime_speaker_embedding_ecapa_tdnn_groupnorm
pipeline_tag: audio-classification
tags:
  - speaker-verification
  - character-identity
  - cross-lingual
  - anime
  - speech
language:
  - ja
  - en

KyaraEmbed

An encoder for anime character identity rather than speaker identity. Given an audio clip, it returns a 192-dimensional embedding whose cosine similarity follows the character across a change of performer, and separates characters who share one.

Fine-tuned from a public anime-domain ECAPA-TDNN. Evaluation numbers and the benchmark this was built for live in the KyaraBench repository.

Checkpoint

File Size Notes
kyaraembed_v1.0.pt 83 MB ECAPA-TDNN, 20.8M parameters, 192-d output
kyaraembed_v1.0.pt.sha256 — Checksum
LICENSE-MODEL, NOTICE — Licence and the upstream attribution chain

How to use

pip install torch soundfile librosa anime-speaker-embedding
import numpy as np, torch, soundfile as sf, librosa
from anime_speaker_embedding import AnimeSpeakerEmbedding

device = "cuda" if torch.cuda.is_available() else "cpu"
model = AnimeSpeakerEmbedding(device=device, variant="char")
model.load_state_dict(torch.load("kyaraembed_v1.0.pt", map_location=device))
model.eval()

def load(path, seconds=10.0, sr=16000):
    """Mono 16 kHz at a fixed length: longer is truncated, shorter is zero-padded."""
    wav, file_sr = sf.read(path, dtype="float32", always_2d=True)
    wav = wav.mean(axis=1)
    if file_sr != sr:
        wav = librosa.resample(wav, orig_sr=file_sr, target_sr=sr)
    out = np.zeros(int(seconds * sr), dtype="float32")
    wav = wav[:len(out)]
    out[:len(wav)] = wav
    return out

x = torch.from_numpy(np.stack([load("a.wav"), load("b.wav")])).to(device)
with torch.no_grad():
    z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
print(f"similarity: {float(z[0] @ z[1]):.3f}")

Keep the fixed length. The encoder pools over time, so padding a clip to whatever else happens to be in its batch changes its embedding; padding to a fixed length does not.

For several clips per side, mean-pool the normalised embeddings and renormalise before the cosine:

def prototype(paths):
    x = torch.from_numpy(np.stack([load(p) for p in paths])).to(device)
    with torch.no_grad():
        z = torch.nn.functional.normalize(model(x).squeeze(1), dim=-1)
    return torch.nn.functional.normalize(z.mean(0), dim=-1)

Licence

CC BY-NC 4.0. Part of the audio behind these weights is held under terms permitting non-commercial research only. See LICENSE-MODEL and NOTICE.