haan: Smart Turn v3.2 fine-tuned for phone audio
End-of-turn detection for voice agents on phone calls. haan is Smart Turn v3.2 fine-tuned with telephone-channel augmentation (8 kHz G.711 mu-law and the 300โ3400 Hz PSTN voice band), so it cuts callers off less often on phone lines.
Code, training and evaluation: https://github.com/Charansripadi/haan
Files
| File | What it is |
|---|---|
haan-telephony.onnx |
FP32 ONNX. Same interface as smart-turn-v3.2-gpu.onnx: input input_features (batch, 80, 800), output logits (batch, 1) holding P(turn complete) after the sigmoid. Drop-in for any Smart Turn v3.2 runner. |
haan-telephony.pt |
PyTorch state dict for haan.model.SmartTurnModel. |
train_meta.json |
Training settings and the 12 training shards used. |
Preprocessing is identical to Smart Turn v3.2: the last 8 s of 16 kHz mono
audio, left-padded with zeros, Whisper feature extractor with do_normalize=True.
Results
Full official Smart Turn v3.2 test set (31,527 clips, 23 languages), threshold 0.5. "control" is the same fine-tune without phone augmentation (same init, data, steps, learning rate and seed); only the channel mix differs.
| control | haan (telephony) | |
|---|---|---|
| Accuracy, PSTN phone line | 91.52% | 92.78% |
| Cut-off rate (FPR), PSTN phone line | 11.09% | 8.64% |
| Accuracy, normal 16 kHz audio | 93.53% | 93.63% |
Phone-line accuracy improved in 23 of 23 languages. On normal audio, accuracy is unchanged overall, with mixed per-language changes (14 up, 7 down).
LiveKit eot-bench
eot-bench scores turn detectors on real human-to-agent conversations in 14 languages, at every pause, under latency and false-cutoff budgets (lower is better). English:
| Cutoffs @ 300 ms | Cutoffs @ 600 ms | Latency @ 5% cutoffs | Latency @ 10% cutoffs | |
|---|---|---|---|---|
| SmartTurn v3.2 | 35.2% | 14.8% | 1051 ms | 739 ms |
| haan | 35.3% | 14.8% | 1044 ms | 742 ms |
On the benchmark's own audio, haan and SmartTurn v3.2 are essentially tied. Across the 14 languages, the mean cutoff rate at 600 ms is 12.9% for haan and 13.3% for SmartTurn v3.2 (haan lower in 10 of 14).
The same audio passed through the PSTN phone-channel simulation (our addition, not part of the official benchmark):
| English, phone line | Cutoffs @ 300 ms | Cutoffs @ 600 ms | Latency @ 5% cutoffs | Latency @ 10% cutoffs |
|---|---|---|---|---|
| SmartTurn v3.2 | 37.7% | 15.6% | 1059 ms | 743 ms |
| haan | 37.0% | 14.5% | 1037 ms | 723 ms |
Across 14 languages on the phone line, the mean cutoff rate at 600 ms is 12.9% for haan and 13.7% for SmartTurn v3.2 (haan lower in 11 of 14), and mean latency at a 10% cutoff budget is 672 ms vs 688 ms (haan faster in 12 of 14). SmartTurn v3.2 degrades on phone audio (13.3% to 13.7%) while haan stays at 12.9%. Differences are small and come without confidence intervals; each language has roughly 760 to 1,150 pauses.
Training
One epoch on 12 of 83 Smart Turn v3.2 training shards (~39k clips), AdamW at lr 1e-5 with cosine decay and 10% warmup, batch 64, fp16, 613 steps on a T4. Each clip went through wideband, narrowband (8 kHz mu-law) or PSTN with equal probability, with a random level of โ12 to +3 dB on phone channels. The starting weights were recovered from Daily's published ONNX file (match to 6e-8).
Limitations
- The phone channel is simulated, not recorded on real calls. Packet loss, echo and codecs other than G.711 are not modelled.
- Evaluated on Smart Turn's own test set, which is clip-level. Gains on benchmarks with different audio and protocols may be smaller.
- Fine-tuned for one epoch on a subset of the training data.
Licence
BSD-2-Clause, following Smart Turn (Copyright Daily).
Model tree for CharanSripadi/haan
Base model
pipecat-ai/smart-turn-v3