turn-1

A server-class model that knows when you have finished talking. It hears the last 8 seconds of the user and reads the agent's last sentence. Built on open speech encoders. Weights and code: Apache-2.0.

On LiveKit's public eot-bench (English), on point estimates:

Model False cut-offs @ 300 ms @ 600 ms Latency @ 5% cut-offs @ 10%
LiveKit Turn Detector v1 9.9% 4.5% 543 ms 295 ms
turn-1 10.6% 4.5% 553 ms 322 ms
JoinIn AI Baton 12.3% 4.8% 577 ms 350 ms
Deepgram Flux 12.9% 9.9% 1,151 ms 548 ms

Lower is better. Intervals and limits are below. For a model that runs on the device, see turn-1-mini (5.3M parameters).

Quick start

pip install "p99turn1 @ https://huggingface.co/p99lab/turn-1/resolve/main/dist/p99turn1-1.0.0-py3-none-any.whl"
pip install --no-deps qwen-asr==0.0.6        # the model code of the speech encoders
from p99turn1 import Turn1

model = Turn1()                               # downloads p99lab/turn-1 and the two speech encoders (4.7 GB + 1.9 GB)
p = model.score(audio_so_far, agent_text="What's your account number?")   # 16 kHz mono float32 -> P(end of turn)

stream = model.stream(agent_text="What's your account number?")
stream.push(chunk)                            # as audio arrives
p = stream.score()                            # ask whenever you need a decision; calibrate the threshold on your own audio

A CUDA GPU with 20 GB is what we ran it on. Made by p99lab (p99lab.com). Version 1.0.0, October 2026. Questions and issues: research@p99lab.com

What it is

turn-1 is the mean of two estimates of P(end of turn), made from the same audio:

Part Whose Size
An open speech encoder, used unchanged Third-party, Apache-2.0. Downloaded from its own repository at a pinned revision; not redistributed here (see NOTICE) 4.7 GB
turn-1-head.safetensors: an end-of-turn head on that model's features p99lab 18.9M parameters, 76 MB
A second, smaller open speech encoder, base weights unchanged Third-party, Apache-2.0. Downloaded from its own repository at a pinned revision; not redistributed here (see NOTICE) 1.9 GB
turn-1-lora.safetensors: a low-rank adapter for that model, added at run time, and a head p99lab 14.8M parameters, 59 MB
Input The last 8 s of 16 kHz mono audio heard so far (low-passed at 4 kHz inside the package) and, optionally, the agent's previous utterance as text
Never used The words of the user's current turn as text: the model hears them
Output P(end of turn). No fixed decision point: ask at any time during a pause
Cost Each of the two models runs once per decision (no text is generated). Not incremental
Precision float32

Evaluation: eot-bench English

livekit/eot-bench-data, config en, split validation, revision ca9d98a, harness commit 9ee21b5: 400 turns, 705 scored mid-turn pauses. float32, batch size 64. Lower is better.

False cut-offs @ 300 ms False cut-offs @ 600 ms Latency @ 5% Latency @ 10% Score AUC (diagnostic)
turn-1 10.6% (75 of 705) 4.5% 553 ms 322 ms 0.976
95% interval (turn-level bootstrap, 400 resamples) 7.7 to 13.8% 2.7 to 6.0% 399 to 691 ms 250 to 387 ms
LiveKit Turn Detector v1 9.9% (70 of 705) 4.5% 543 ms 295 ms 0.969
JoinIn AI Baton 12.3% (87 of 705) 4.8% 577 ms 350 ms 0.958

What this supports, and what it does not:

  • Counted against the 13 rows of the board at that commit, turn-1 would be second on false cut-offs at 300 ms, on latency at 5% and on latency at 10%, and tied first on false cut-offs at 600 ms (32 of 705, the same count as LiveKit Turn Detector v1).
  • It is behind LiveKit Turn Detector v1 at 300 ms, at 5% and at 10%.
  • It is ahead of JoinIn AI Baton on all four point estimates. No paired test against any other entry was made, and the single-model intervals overlap.
  • Two candidates of turn-1 were scored on this test set. The first scored 12.9% / 4.8% / 593 ms / 363 ms. This one was kept under a rule fixed before it was scored. Paired on the same turns, the difference between the two includes zero on all four measures (at 300 ms: -2.3 points, 95% interval -4.5 to +0.4).
  • English only. No other language was run and no multilingual number is claimed.

Latency here is dead air chosen by the benchmark's policy sweep, not compute time. As for every model on the board, the sweep picks threshold, action delay and timeout on the same 400 turns it scores.

eot-bench audio and labels were never used to train, select or threshold this model. The package in this repository reproduces the numbers above.

Limits

  • English only measured.
  • Heavy. Two speech models (2.0B and 0.8B parameters) run for every decision; this is a GPU server model. Time per decision has not been measured in this configuration.
  • 4 kHz band. All input is low-passed at 4 kHz.
  • Scores are not calibrated. Calibrate thresholds on your own audio.
  • Weakest on pauses inside numbers, letters and email addresses (13 of 126 cut off at its 300 ms operating point; LiveKit Turn Detector v1: 8) and after a finished sentence when more is coming (25 of 119; LiveKit v1: 16).
  • No accent, age or gender breakdown; not evaluated with overlapping speech, far-field audio or agent echo.

Responsible use

turn-1 outputs one number about turn-taking. It does not identify speakers, transcribe, or infer emotion, health or identity. A false "end of turn" interrupts a person and a missed one makes them wait: keep a timeout fallback and do not rely on it alone where either could cause harm. Performance for accents, speech impairments and non-English speakers is unmeasured; test on your own users.

Licence

p99lab's weights and code: Apache-2.0. The two speech encoders stay under their own licence (Apache-2.0) and are obtained from their authors' repositories. Third-party credits: see NOTICE.

Citation

@misc{p99lab2026turn1,
  title  = {turn-1: an end-of-turn detection model},
  author = {p99lab},
  year   = {2026},
  url    = {https://p99lab.com}
}

Please also cite eot-bench (LiveKit) when reporting numbers.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results