haan: Smart Turn v3.2 fine-tuned for phone audio

End-of-turn detection for voice agents on phone calls. haan is Smart Turn v3.2 fine-tuned with telephone-channel augmentation (8 kHz G.711 mu-law and the 300โ€“3400 Hz PSTN voice band), so it cuts callers off less often on phone lines.

Code, training and evaluation: https://github.com/Charansripadi/haan

Files

File What it is
haan-telephony.onnx FP32 ONNX. Same interface as smart-turn-v3.2-gpu.onnx: input input_features (batch, 80, 800), output logits (batch, 1) holding P(turn complete) after the sigmoid. Drop-in for any Smart Turn v3.2 runner.
haan-telephony.pt PyTorch state dict for haan.model.SmartTurnModel.
train_meta.json Training settings and the 12 training shards used.

Preprocessing is identical to Smart Turn v3.2: the last 8 s of 16 kHz mono audio, left-padded with zeros, Whisper feature extractor with do_normalize=True.

Results

Full official Smart Turn v3.2 test set (31,527 clips, 23 languages), threshold 0.5. "control" is the same fine-tune without phone augmentation (same init, data, steps, learning rate and seed); only the channel mix differs.

control haan (telephony)
Accuracy, PSTN phone line 91.52% 92.78%
Cut-off rate (FPR), PSTN phone line 11.09% 8.64%
Accuracy, normal 16 kHz audio 93.53% 93.63%

Phone-line accuracy improved in 23 of 23 languages. On normal audio, accuracy is unchanged overall, with mixed per-language changes (14 up, 7 down).

LiveKit eot-bench

eot-bench scores turn detectors on real human-to-agent conversations in 14 languages, at every pause, under latency and false-cutoff budgets (lower is better). English:

Cutoffs @ 300 ms Cutoffs @ 600 ms Latency @ 5% cutoffs Latency @ 10% cutoffs
SmartTurn v3.2 35.2% 14.8% 1051 ms 739 ms
haan 35.3% 14.8% 1044 ms 742 ms

On the benchmark's own audio, haan and SmartTurn v3.2 are essentially tied. Across the 14 languages, the mean cutoff rate at 600 ms is 12.9% for haan and 13.3% for SmartTurn v3.2 (haan lower in 10 of 14).

The same audio passed through the PSTN phone-channel simulation (our addition, not part of the official benchmark):

English, phone line Cutoffs @ 300 ms Cutoffs @ 600 ms Latency @ 5% cutoffs Latency @ 10% cutoffs
SmartTurn v3.2 37.7% 15.6% 1059 ms 743 ms
haan 37.0% 14.5% 1037 ms 723 ms

Across 14 languages on the phone line, the mean cutoff rate at 600 ms is 12.9% for haan and 13.7% for SmartTurn v3.2 (haan lower in 11 of 14), and mean latency at a 10% cutoff budget is 672 ms vs 688 ms (haan faster in 12 of 14). SmartTurn v3.2 degrades on phone audio (13.3% to 13.7%) while haan stays at 12.9%. Differences are small and come without confidence intervals; each language has roughly 760 to 1,150 pauses.

Training

One epoch on 12 of 83 Smart Turn v3.2 training shards (~39k clips), AdamW at lr 1e-5 with cosine decay and 10% warmup, batch 64, fp16, 613 steps on a T4. Each clip went through wideband, narrowband (8 kHz mu-law) or PSTN with equal probability, with a random level of โˆ’12 to +3 dB on phone channels. The starting weights were recovered from Daily's published ONNX file (match to 6e-8).

Limitations

  • The phone channel is simulated, not recorded on real calls. Packet loss, echo and codecs other than G.711 are not modelled.
  • Evaluated on Smart Turn's own test set, which is clip-level. Gains on benchmarks with different audio and protocols may be smaller.
  • Fine-tuned for one epoch on a subset of the training data.

Licence

BSD-2-Clause, following Smart Turn (Copyright Daily).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for CharanSripadi/haan

Quantized
(6)
this model