ReDimNet2-B6 Core ML Speaker Embeddings

ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.

Model

Property Value
Parameters 12.3 million
Format Compiled Core ML; FP32 frontend/head, FP16 backbone
Compiled size 28.6 MiB
Input 96,000 mono Float32 samples
Sample rate 16 kHz
Window 6 seconds
Output 192-dimensional L2-normalized embedding
Minimum deployment macOS 15 / iOS 18

The checkpoint was trained on VoxBlink2 and VoxCeleb2. The fixed six-second shape avoids the slow Core ML fallback observed with a flexible waveform shape. Applications should repeat clean two-to-six-second speech to fill the input and center-crop longer samples.

Export revision frontend-fp32-v1 preserves waveform normalization, spectral and mel computation, log/feature normalization, pooling and output projection in FP32. The learned backbone remains FP16. This prevents numerical overflow in preprocessing without changing the checkpoint or input/output contract.

Files

File Size Description
ReDimNet2B6.mlmodelc/ 28.6 MiB Precompiled Core ML model
config.json <16 KiB Contract, export revision, compiled-file SHA-256 and numerical validation
checksums.json <2 KiB SHA-256 of every other published artifact
README.md <8 KiB This model card
LICENSE 1.0 KiB MIT license from the upstream implementation

Performance

Measured on an Apple M5 Pro after two warm-up predictions:

Measurement Result Meaning
Warm six-second inference 13.6 ms One voice-profile embedding
Warm throughput 73.6 embeddings/s Repeated six-second windows after warm-up
Numerical validation 20/20 configurations and controls Five deterministic controls across four allowed-device settings

Numerical validation checks finite, normalized outputs and cosine >= 0.999 against the original FP32 model. Controls cover a noisy two-tone waveform, sparse burst, quiet waveform, repeated 600 ms waveform and silence. Allowed-device settings do not identify actual per-operation placement. These checks do not measure diarization, recognition accuracy, or speaker-verification error rate. Thresholds must be calibrated for the intended microphones, languages, and acoustic conditions. Speaker embeddings are useful for labeling; they are not biometric authentication and do not protect against voice spoofing.

Python usage

import coremltools as ct
import numpy as np

model = ct.models.CompiledMLModel("ReDimNet2B6.mlmodelc")
audio = np.zeros((1, 96_000), dtype=np.float32)
embedding = model.predict({"audio": audio})["embedding"]

speech-swift

speech embed-speaker voice.wav --engine redimnet2 --json
import SpeechVAD

let model = try await ReDimNet2SpeakerModel.fromPretrained()
let embedding = try model.embed(audio: samples, sampleRate: 16_000)

Source

Converted from the official PalabraAI/ReDimNet2 B6 vb2+vox2_v0 large-margin checkpoint. The source revision and checkpoint SHA-256 are recorded in config.json.

Links

Downloads last month
262
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including aufklarer/ReDimNet2-B6-CoreML