Papers
arxiv:2609.23114

Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations

Published on Sep 19
Authors:
,
,

Abstract

Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) and evaluated using word-level metrics, making it difficult to assess SD performance independent of ASR accuracy. In this work, we investigate ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input. We systematically compare two output representations: an event-based representation that explicitly models speaker turn onset and offset timestamps, and a frame-based representation that predicts frame-level speaker activity. To provide structured conversational cues, we further incorporate auxiliary tasks including speech activity detection, overlapped speech detection, and speaker turn counting within the output sequence. Across multiple meeting datasets, we find that event-based representations produce more stable and consistent SD outputs than frame-based representations. Our analysis shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recording-level speaker tracking and overlap-related misses. Explicit speaker-linking post-processing substantially reduces speaker confusion, suggesting that robust SpeechLM-based SD requires persistent speaker tracking and overlap-aware generation.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.23114
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.23114 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.23114 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.23114 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.