Instructions to use IOTEverythin/cue-v5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IOTEverythin/cue-v5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="IOTEverythin/cue-v5")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("IOTEverythin/cue-v5", device_map="auto") - Notebooks
- Google Colab
- Kaggle

Cue v5: turn-taking for voice agents, trained on real conversation
Paper: Cue: A Conversation Reflex Model for Barge-In and End-of-Turn Control in Cascaded Voice Agents, preprint, 2026. doi:10.5281/zenodo.23144482 · OpenAIRE research archive
Cue listens to a call and decides, every 160 ms, what a voice agent should do: keep talking, PAUSE or STOP while it speaks; RESPOND or keep waiting after the caller stops. v5 is Cue v4 retrained on real human conversation as well as synthetic calls, and it can hear the agent's own audio.
What changed from v4
- Real conversation. Besides the synthetic Indian English calls, v5 was trained on the AMI and ICSI meeting corpora (CC BY 4.0): real backchannels, overlaps, interruptions and hesitations. On held-out real meetings it catches about twice as many interruptions as v4.
- The agent's own audio as a second input. A live agent knows what it is playing; v5 hears it through a second encoder pass, which helps it tell the agent's own voice (echo) from the caller. It is optional: without it the input is zero, a condition v5 was also trained on.
- Phone conditions. Meetings and calls were trained clean and through echo, reverb, recorded noise (MUSAN) and an 8 kHz phone codec.
- No shortcut through the agent's reply. In recordings the other side starts talking right after a turn ends; a live agent only does after Cue decides. The agent audio is hidden in that window and meetings are trained without the other side's audio, so v5 does not learn that shortcut.
Status: early release. Real meetings are not phone calls, and v5 has not been tested on production call traffic.
Use
import cue_turn # pip install "cue-turn[full]"
cue = cue_turn.load("v5") # or a folder with this repository's files
stream = cue.stream(sample_rate=8000, agent_sample_rate=24000)
for caller_chunk, agent_chunk in call: # the agent chunk: what it played over the same span
for d in stream.feed(caller_chunk, assistant_speaking=bot_is_playing(), agent_audio=agent_chunk):
handle(d) # {"t_ms": ..., "decision": "STOP" | "PAUSE" | "RESPOND"}
In Pipecat, add CueAgentAudioTap(session) just before transport.output() so the session sees the
bot's audio; see the cue-turn README. A GPU is recommended (two encoder passes while the agent
speaks; one while it is quiet).
Decision settings
| Profile | Barge-in (stop_p, hold) | End of turn (p_done, turn_hold) |
|---|---|---|
| responsive | 0.7, 4 | 0.5, 6 |
| balanced (default) | 0.8, 6 | 0.9, 12 |
| cautious | 0.95, 6 | 0.95, 16 |
All with cooldown 1500 ms, min_speech 10 frames, fallback 2000 ms. Defaults were chosen on AMI validation meetings and the synthetic test sets together: real speech wants lower thresholds for interruptions and longer holds for end of turn than synthetic speech does.
| Setting | Synthetic barge-in, clean: hard / takeover / stops on backchannel | AMI validation: interruptions caught / false alarms |
|---|---|---|
| 0.7, 4 | 85.5% / 96.9% / 11.4% | 63% / 7.8% |
| 0.8, 6 | 90.3% / 93.8% / 8.8% | 39% / 4.1% |
| 0.95, 6 | 91.9% / 87.5% / 3.5% | 2% / 0% |
| Setting | Synthetic end of turn, clean: mid-thought / answered / delay | AMI validation: turn ends caught / answers into pauses |
|---|---|---|
| 0.5, 6 | 4.5% / 95.6% / 137 ms | 76% / 37% |
| 0.9, 12 | 3.7% / 96.7% / 256 ms | 63% / 25% |
| 0.95, 16 | 3.0% / 96.6% / 336 ms | 51% / 14% |
Results
Real meetings (held-out AMI and ICSI test meetings; best setting with at most 10% false alarms)
| v4 | v5 | |
|---|---|---|
| AMI interruptions caught | 24% | 61% |
| ICSI interruptions caught | 37% | 63% |
| AMI end of turn | none under 10% false alarms | 27% (1.2 s) |
| ICSI end of turn | none under 10% false alarms | 59% (1.06 s) |
Synthetic test calls (each model's best setting under the v4 card's matched rules)
| Hard | Takeovers | Soft | Stops on backchannel | Stops on noise | Reaction p50 | |
|---|---|---|---|---|---|---|
| v4, clean | 85.5% | 90.6% | 93.8% | 3.5% | 5.6% | 243 ms |
| v5, clean | 90.3% | 93.8% | 92.5% | 7.0% | 7.4% | 241 ms |
| v4, realistic | 75.8% | 75.0% | 91.2% | 3.5% | 8.0% | 311 ms |
| v5, realistic | 85.5% | 81.2% | 93.8% | 4.4% | 11.7% | 281 ms |
| End of turn | Answers mid-thought | Turn ends answered | Delay p50 |
|---|---|---|---|
| v4, clean / realistic | 5.2% / 7.5% | 95.5% / 94.4% | 177 / 185 ms |
| v5, clean / realistic | 4.5% / 7.5% | 95.6% / 94.1% | 137 / 154 ms |
TurnBench (dev set, cross-validated settings, official scorer)
Both models without any TurnBench training data. Recall / false alarms / median delay.
| End of turn | End of turn, Cue's end-of-turn judgement | Interruption | |
|---|---|---|---|
| v4 | 0.772 / 7.7% / 1.41 s | 0.608 / 8.7% / 1.16 s | 0.614 / 8.8% / 583 ms |
| v5, without the other speaker's audio | 0.779 / 8.9% / 1.38 s | 0.605 / 7.0% / 947 ms | 0.804 / 7.9% / 547 ms |
| v5, with the other speaker's audio | 0.800 / 7.0% / 1.16 s | 0.686 / 8.1% / 908 ms | 0.712 / 9.2% / 653 ms |
The best end-of-turn rule on TurnBench is a silence timer on Cue's speech detection (about 1 s of quiet): the benchmark ranks recall within a false-alarm budget and does not count delay. Published systems on the same dev set include VAP (0.841 / 0.957) and Kyutai's semantic VAD (0.803 / 0.934).
Training
| Start | Cue v4 (the agent-audio weights start at zero, so step 0 is v4) |
| Data | 2,600 synthetic call tracks (1,300 calls, clean + realistic; 32 with Dia-1.6B caller voices), AMI 20 meetings (80 speaker tracks), ICSI 10 meetings (30 speaker tracks); sampled 50 / 30 / 20% |
| Labels | synthetic: exact event timings from the generator; meetings: from the corpus dialogue acts (backchannels, floor-taking, interrupted utterances), word timings and vocal sounds |
| Augmentation | echo, room reverb, MUSAN noise, 8 kHz mu-law codec, level, on half the tracks |
| Agent input | the agent's audio, kept only while the agent is playing and hidden in the 1.5 s decision window; zeroed for 15% of synthetic segments and for all meetings |
| Optimiser | AdamW, lr 1e-4, one-cycle, batch 16 x 20 s, 8,000 steps, bfloat16 on an RTX 4060 laptop GPU (79 min) |
Limitations
- Real meetings are not phone calls; v5 has not been validated on production calls.
- End of turn on real speech is still slow (about 1 s) and catches about half of turn ends.
- The default settings trade synthetic false stops for real-speech recall; tune per deployment.
- Indian-English and Hinglish backchannels ("haan", "achha") are only in the synthetic data.
Citation
@misc{cue2026,
title = {Cue: A Conversation Reflex Model for Barge-In and End-of-Turn Control in Cascaded Voice Agents},
author = {Nishanth Tarun A, Joshua and Ajitesh Varun A, Joel},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.23144482},
url = {https://doi.org/10.5281/zenodo.23144482},
note = {Preprint}
}
Licences
encoder/: Whisper small's encoder, Apache-2.0 (position table shortened; see NOTICE).cue.safetensors,config.json: IOTEverythin, Apache-2.0 (see LICENSE). Commercial use is allowed; keep the NOTICE file, which carries the attributions the CC BY 4.0 training data requires.- Training data: AMI and ICSI (CC BY 4.0), MUSAN (CC BY 4.0), Chatterbox (MIT), Dia-1.6B (Apache-2.0), GLOBE (CC0), Svarah (CC BY 4.0). See NOTICE.
- Downloads last month
- -
Model tree for IOTEverythin/cue-v5
Base model
openai/whisper-small