Instructions to use pyannote/segmentation-3.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- pyannote.audio
How to use pyannote/segmentation-3.0 with pyannote.audio:
from pyannote.audio import Model, Inference model = Model.from_pretrained("pyannote/segmentation-3.0") inference = Inference(model) # inference on the whole file inference("file.wav") # inference on an excerpt from pyannote.core import Segment excerpt = Segment(start=2.0, end=5.0) inference.crop("file.wav", excerpt) - Notebooks
- Google Colab
- Kaggle
Question: edge/mobile deployment β anyone tested?
We benchmark models on 40 phones (Snapdragon 865) at Dispatch AI (FZE, UAE).
Question: has anyone tested this model on mobile/edge? Interested in:
- Inference speed (t/s)
- Model size after quantization
- RAM usage
Happy to share phone farm benchmark results.
- Dispatch AI (FZE), Sharjah UAE
Worth knowing there is a CoreML port at FluidInference/speaker-diarization-coreml. Apple-only, so not directly useful for a Snapdragon farm, but it does show the pipeline moves off the server without much trouble.
One thing that tends to surprise people on device: segmentation-3.0 is small enough that quantized size is rarely the constraint. Cost is dominated by how many overlapping windows you run over the audio. And in the full diarization pipeline the speaker embedding step is usually heavier than segmentation itself, so those two are worth measuring separately rather than as one number.