ReDimNet2-B6 Core ML Speaker Embeddings
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.
Model
| Property | Value |
|---|---|
| Parameters | 12.3 million |
| Format | Compiled Core ML; FP32 frontend/head, FP16 backbone |
| Compiled size | 28.6 MiB |
| Input | 96,000 mono Float32 samples |
| Sample rate | 16 kHz |
| Window | 6 seconds |
| Output | 192-dimensional L2-normalized embedding |
| Minimum deployment | macOS 15 / iOS 18 |
The checkpoint was trained on VoxBlink2 and VoxCeleb2. The fixed six-second shape avoids the slow Core ML fallback observed with a flexible waveform shape. Applications should repeat clean two-to-six-second speech to fill the input and center-crop longer samples.
Export revision frontend-fp32-v1 preserves waveform normalization, spectral
and mel computation, log/feature normalization, pooling and output projection
in FP32. The learned backbone remains FP16. This prevents numerical overflow
in preprocessing without changing the checkpoint or input/output contract.
Files
| File | Size | Description |
|---|---|---|
ReDimNet2B6.mlmodelc/ |
28.6 MiB | Precompiled Core ML model |
config.json |
<16 KiB | Contract, export revision, compiled-file SHA-256 and numerical validation |
checksums.json |
<2 KiB | SHA-256 of every other published artifact |
README.md |
<8 KiB | This model card |
LICENSE |
1.0 KiB | MIT license from the upstream implementation |
Performance
Measured on an Apple M5 Pro after two warm-up predictions:
| Measurement | Result | Meaning |
|---|---|---|
| Warm six-second inference | 13.6 ms | One voice-profile embedding |
| Warm throughput | 73.6 embeddings/s | Repeated six-second windows after warm-up |
| Numerical validation | 20/20 configurations and controls | Five deterministic controls across four allowed-device settings |
Numerical validation checks finite, normalized outputs and cosine >= 0.999 against the original FP32 model. Controls cover a noisy two-tone waveform, sparse burst, quiet waveform, repeated 600 ms waveform and silence. Allowed-device settings do not identify actual per-operation placement. These checks do not measure diarization, recognition accuracy, or speaker-verification error rate. Thresholds must be calibrated for the intended microphones, languages, and acoustic conditions. Speaker embeddings are useful for labeling; they are not biometric authentication and do not protect against voice spoofing.
Python usage
import coremltools as ct
import numpy as np
model = ct.models.CompiledMLModel("ReDimNet2B6.mlmodelc")
audio = np.zeros((1, 96_000), dtype=np.float32)
embedding = model.predict({"audio": audio})["embedding"]
speech-swift
speech embed-speaker voice.wav --engine redimnet2 --json
import SpeechVAD
let model = try await ReDimNet2SpeakerModel.fromPretrained()
let embedding = try model.embed(audio: samples, sampleRate: 16_000)
Source
Converted from the official
PalabraAI/ReDimNet2 B6
vb2+vox2_v0 large-margin checkpoint. The source revision and checkpoint
SHA-256 are recorded in config.json.
Links
- speech-swift — Apple SDK
- Docs — install and CLI docs
- soniqo.audio
- blog
- Downloads last month
- 262