File size: 2,546 Bytes
05b034e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
license: apache-2.0
base_model: LocalAI-io/LocalVQE
tags:
  - coreml
  - audio
  - speech-enhancement
  - acoustic-echo-cancellation
  - noise-suppression
  - fluidaudio
library_name: fluidaudio
---

# LocalVQE Core ML

Streaming Core ML exports of [LocalVQE](https://github.com/localai-org/LocalVQE)
(weights: [LocalAI-io/LocalVQE](https://huggingface.co/LocalAI-io/LocalVQE),
Apache-2.0): neural acoustic echo cancellation + noise suppression +
dereverberation for 16 kHz speech, a CPU-tuned derivative of DeepVQE
(Indenbom et al., Interspeech 2023).

Consumed by [FluidAudio](https://github.com/FluidInference/FluidAudio)
(`LocalVqeManager` / `LocalVqeStream`); conversion code in
[FluidInference/mobius](https://github.com/FluidInference/mobius)
`models/enhancement/localvqe/coreml`.

## Files

| File | Checkpoint | Params | Samples per call |
|---|---|---:|---:|
| `localvqe-v1.3-4.8M-256ms.mlmodelc` | `localvqe-v1.3-4.8M.pt` | 4.8 M | 4096 (16 hops) |
| `localvqe-v1.3-4.8M-16ms.mlmodelc` | `localvqe-v1.3-4.8M.pt` | 4.8 M | 256 (1 hop) |
| `localvqe-v1.2-1.3M-256ms.mlmodelc` | `localvqe-v1.2-1.3M.pt` | 1.3 M | 4096 (16 hops) |
| `localvqe-v1.2-1.3M-16ms.mlmodelc` | `localvqe-v1.2-1.3M.pt` | 1.3 M | 256 (1 hop) |

All fp32, iOS 17 / macOS 14 minimum deployment target. The two chunk sizes
produce identical audio; they trade per-call overhead against latency.

## Model I/O

Inputs (Float32): `mic` `[1, N]`, `ref` `[1, N]` (far-end reference — what the
loudspeaker played), and 33 `in_<state>` tensors. Outputs: `enhanced` `[1, N]`
and the matching `out_<state>` tensors. Start with all states zero and feed
each call's `out_*` back as the next call's `in_*`.

The enhanced hop lags the input by 256 samples (16 ms): after consuming input
hop *k* the model emits input hop *k−1*. The first emitted hop covers t < 0
and can be dropped; feed one hop of zeros at the end to drain. Output level
matches the upstream GGML engine (2× the upstream PyTorch reference's
overlap-add convention).

## Parity / speed

Swift output vs the upstream GGML CLI on the upstream double-talk demo clip:
2.8e-5 max abs diff, 80 dB SNR. Apple M5 Pro, per-call p50 on CPU: v1.3 256 ms
7.1 ms (36× RT), v1.3 16 ms 1.2 ms (14× RT), v1.2 256 ms 4.2 ms (60× RT),
v1.2 16 ms 0.7 ms (24× RT).

## Citation

Cite the upstream repository (`CITATION.cff` in
[localai-org/LocalVQE](https://github.com/localai-org/LocalVQE)) and the
DeepVQE paper it derives from (Indenbom et al., Interspeech 2023,
[arXiv:2306.03177](https://arxiv.org/abs/2306.03177)).