Question: edge/mobile deployment β€” anyone tested?

#11
by 3morixd - opened

We benchmark models on 40 phones (Snapdragon 865) at Dispatch AI (FZE, UAE).

Question: has anyone tested this model on mobile/edge? Interested in:

  • Inference speed (t/s)
  • Model size after quantization
  • RAM usage

Happy to share phone farm benchmark results.

  • Dispatch AI (FZE), Sharjah UAE

Worth knowing there is a CoreML port at FluidInference/speaker-diarization-coreml. Apple-only, so not directly useful for a Snapdragon farm, but it does show the pipeline moves off the server without much trouble.

One thing that tends to surprise people on device: segmentation-3.0 is small enough that quantized size is rarely the constraint. Cost is dominated by how many overlapping windows you run over the audio. And in the full diarization pipeline the speaker embedding step is usually heavier than segmentation itself, so those two are worth measuring separately rather than as one number.

Sign up or log in to comment