Cue v5: caller audio and the agent's own audio in; STOP, PAUSE or RESPOND out. 2.5x more real interruptions caught than Cue v4 on AMI meetings; 0.80 interruption recall on TurnBench.

Cue v5: turn-taking for voice agents, trained on real conversation

DOI

Paper: Cue: A Conversation Reflex Model for Barge-In and End-of-Turn Control in Cascaded Voice Agents, preprint, 2026. doi:10.5281/zenodo.23144482 · OpenAIRE research archive

Cue listens to a call and decides, every 160 ms, what a voice agent should do: keep talking, PAUSE or STOP while it speaks; RESPOND or keep waiting after the caller stops. v5 is Cue v4 retrained on real human conversation as well as synthetic calls, and it can hear the agent's own audio.

What changed from v4

  • Real conversation. Besides the synthetic Indian English calls, v5 was trained on the AMI and ICSI meeting corpora (CC BY 4.0): real backchannels, overlaps, interruptions and hesitations. On held-out real meetings it catches about twice as many interruptions as v4.
  • The agent's own audio as a second input. A live agent knows what it is playing; v5 hears it through a second encoder pass, which helps it tell the agent's own voice (echo) from the caller. It is optional: without it the input is zero, a condition v5 was also trained on.
  • Phone conditions. Meetings and calls were trained clean and through echo, reverb, recorded noise (MUSAN) and an 8 kHz phone codec.
  • No shortcut through the agent's reply. In recordings the other side starts talking right after a turn ends; a live agent only does after Cue decides. The agent audio is hidden in that window and meetings are trained without the other side's audio, so v5 does not learn that shortcut.

Status: early release. Real meetings are not phone calls, and v5 has not been tested on production call traffic.

Use

import cue_turn                                   # pip install "cue-turn[full]"
cue = cue_turn.load("v5")                         # or a folder with this repository's files
stream = cue.stream(sample_rate=8000, agent_sample_rate=24000)
for caller_chunk, agent_chunk in call:            # the agent chunk: what it played over the same span
    for d in stream.feed(caller_chunk, assistant_speaking=bot_is_playing(), agent_audio=agent_chunk):
        handle(d)                                 # {"t_ms": ..., "decision": "STOP" | "PAUSE" | "RESPOND"}

In Pipecat, add CueAgentAudioTap(session) just before transport.output() so the session sees the bot's audio; see the cue-turn README. A GPU is recommended (two encoder passes while the agent speaks; one while it is quiet).

Decision settings

Profile Barge-in (stop_p, hold) End of turn (p_done, turn_hold)
responsive 0.7, 4 0.5, 6
balanced (default) 0.8, 6 0.9, 12
cautious 0.95, 6 0.95, 16

All with cooldown 1500 ms, min_speech 10 frames, fallback 2000 ms. Defaults were chosen on AMI validation meetings and the synthetic test sets together: real speech wants lower thresholds for interruptions and longer holds for end of turn than synthetic speech does.

Setting Synthetic barge-in, clean: hard / takeover / stops on backchannel AMI validation: interruptions caught / false alarms
0.7, 4 85.5% / 96.9% / 11.4% 63% / 7.8%
0.8, 6 90.3% / 93.8% / 8.8% 39% / 4.1%
0.95, 6 91.9% / 87.5% / 3.5% 2% / 0%
Setting Synthetic end of turn, clean: mid-thought / answered / delay AMI validation: turn ends caught / answers into pauses
0.5, 6 4.5% / 95.6% / 137 ms 76% / 37%
0.9, 12 3.7% / 96.7% / 256 ms 63% / 25%
0.95, 16 3.0% / 96.6% / 336 ms 51% / 14%

Results

Real meetings (held-out AMI and ICSI test meetings; best setting with at most 10% false alarms)

v4 v5
AMI interruptions caught 24% 61%
ICSI interruptions caught 37% 63%
AMI end of turn none under 10% false alarms 27% (1.2 s)
ICSI end of turn none under 10% false alarms 59% (1.06 s)

Synthetic test calls (each model's best setting under the v4 card's matched rules)

Hard Takeovers Soft Stops on backchannel Stops on noise Reaction p50
v4, clean 85.5% 90.6% 93.8% 3.5% 5.6% 243 ms
v5, clean 90.3% 93.8% 92.5% 7.0% 7.4% 241 ms
v4, realistic 75.8% 75.0% 91.2% 3.5% 8.0% 311 ms
v5, realistic 85.5% 81.2% 93.8% 4.4% 11.7% 281 ms
End of turn Answers mid-thought Turn ends answered Delay p50
v4, clean / realistic 5.2% / 7.5% 95.5% / 94.4% 177 / 185 ms
v5, clean / realistic 4.5% / 7.5% 95.6% / 94.1% 137 / 154 ms

TurnBench (dev set, cross-validated settings, official scorer)

Both models without any TurnBench training data. Recall / false alarms / median delay.

End of turn End of turn, Cue's end-of-turn judgement Interruption
v4 0.772 / 7.7% / 1.41 s 0.608 / 8.7% / 1.16 s 0.614 / 8.8% / 583 ms
v5, without the other speaker's audio 0.779 / 8.9% / 1.38 s 0.605 / 7.0% / 947 ms 0.804 / 7.9% / 547 ms
v5, with the other speaker's audio 0.800 / 7.0% / 1.16 s 0.686 / 8.1% / 908 ms 0.712 / 9.2% / 653 ms

The best end-of-turn rule on TurnBench is a silence timer on Cue's speech detection (about 1 s of quiet): the benchmark ranks recall within a false-alarm budget and does not count delay. Published systems on the same dev set include VAP (0.841 / 0.957) and Kyutai's semantic VAD (0.803 / 0.934).

Training

Start Cue v4 (the agent-audio weights start at zero, so step 0 is v4)
Data 2,600 synthetic call tracks (1,300 calls, clean + realistic; 32 with Dia-1.6B caller voices), AMI 20 meetings (80 speaker tracks), ICSI 10 meetings (30 speaker tracks); sampled 50 / 30 / 20%
Labels synthetic: exact event timings from the generator; meetings: from the corpus dialogue acts (backchannels, floor-taking, interrupted utterances), word timings and vocal sounds
Augmentation echo, room reverb, MUSAN noise, 8 kHz mu-law codec, level, on half the tracks
Agent input the agent's audio, kept only while the agent is playing and hidden in the 1.5 s decision window; zeroed for 15% of synthetic segments and for all meetings
Optimiser AdamW, lr 1e-4, one-cycle, batch 16 x 20 s, 8,000 steps, bfloat16 on an RTX 4060 laptop GPU (79 min)

Limitations

  • Real meetings are not phone calls; v5 has not been validated on production calls.
  • End of turn on real speech is still slow (about 1 s) and catches about half of turn ends.
  • The default settings trade synthetic false stops for real-speech recall; tune per deployment.
  • Indian-English and Hinglish backchannels ("haan", "achha") are only in the synthetic data.

Citation

@misc{cue2026,
  title     = {Cue: A Conversation Reflex Model for Barge-In and End-of-Turn Control in Cascaded Voice Agents},
  author    = {Nishanth Tarun A, Joshua and Ajitesh Varun A, Joel},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.23144482},
  url       = {https://doi.org/10.5281/zenodo.23144482},
  note      = {Preprint}
}

Licences

  • encoder/: Whisper small's encoder, Apache-2.0 (position table shortened; see NOTICE).
  • cue.safetensors, config.json: IOTEverythin, Apache-2.0 (see LICENSE). Commercial use is allowed; keep the NOTICE file, which carries the attributions the CC BY 4.0 training data requires.
  • Training data: AMI and ICSI (CC BY 4.0), MUSAN (CC BY 4.0), Chatterbox (MIT), Dia-1.6B (Apache-2.0), GLOBE (CC0), Svarah (CC BY 4.0). See NOTICE.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IOTEverythin/cue-v5

Finetuned
(3784)
this model