Whisper-Small: Optimized for Qualcomm Devices

HuggingFace Whisper-Small ASR (Automatic Speech Recognition) model is a state-of-the-art system designed for transcribing spoken language into written text. This model is based on the transformer architecture and has been optimized for edge inference by replacing Multi-Head Attention (MHA) with Single-Head Attention (SHA) and linear layers with convolutional (conv) layers. It exhibits robust performance in realistic, noisy environments, making it highly reliable for real-world applications. Specifically, it excels in long-form transcription, capable of accurately transcribing audio clips up to 30 seconds long. Time to the first token is the encoder's latency, while time to each additional token is decoder's latency, where we assume a max decoded length specified below.

This is based on the implementation of Whisper-Small found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.

Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.

Deploying Whisper-Small on-device

This model is compatible with the Qualcomm Voice AI SDK. Download the SDK from the Qualcomm Package Manager to deploy this model on-device.

Getting Started

There are two ways to deploy this model on your device:

Option 1: Download Pre-Exported Models

Below are pre-exported model assets ready for deployment.

Runtime Precision Chipset SDK Versions Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Snapdragon® X Elite QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50, ONNX Runtime 1.27.1 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® X2 Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® X Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA8775P QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA7255P QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA8295P QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® X2 Elite QAIRT 2.50 Download
VOICE_AI float Snapdragon® X Elite QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA8775P QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA7255P QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA8295P QAIRT 2.50 Download

For more device-specific assets and performance metrics, visit Whisper-Small on Qualcomm® AI Hub.

Option 2: Export with Custom Configurations

Use the Qualcomm® AI Hub Models Python library to compile and export the model with your own:

  • Custom weights (e.g., fine-tuned checkpoints)
  • Custom input shapes
  • Target device and runtime configurations

This option is ideal if you need to customize the model beyond the default configuration provided here.

See our repository for Whisper-Small on GitHub for usage instructions.

Model Details

Model Type: Model_use_case.speech_recognition

Model Stats:

  • Input resolution: 80x3000 (30 seconds audio)
  • Max decoded sequence length: 200 tokens
  • Model checkpoint: openai/whisper-small
  • Model size (decoder) (float): 533 MB
  • Model size (encoder) (float): 391 MB
  • Number of parameters (decoder): 139M
  • Number of parameters (encoder): 102M

Performance Summary

Model Runtime Precision Chipset Inference Time (ms) Peak Memory Range (MB) Primary Compute Unit
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.378 ms 44 - 57 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile 8.363 ms 57 - 69 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite 6.279 ms 60 - 60 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® X Elite 10.474 ms 287 - 287 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile 9.933 ms 75 - 87 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile 14.147 ms 74 - 89 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 14.94 ms 60 - 124 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.404 ms 0 - 318 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® QCS8450 14.147 ms 74 - 89 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 13.682 ms 60 - 123 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-X7181 10.474 ms 287 - 287 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ Q-8750 8.363 ms 57 - 69 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.3 ms 45 - 54 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile 8.265 ms 0 - 9 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® X2 Elite 6.773 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® X Elite 11.107 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile 9.788 ms 60 - 68 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile 13.898 ms 60 - 75 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 14.524 ms 60 - 130 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.45 ms 60 - 62 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8775P 13.849 ms 34 - 43 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8650P 13.849 ms 34 - 43 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8255P 13.849 ms 34 - 43 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® QCS8450 13.898 ms 60 - 75 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 13.553 ms 60 - 129 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-X7181 11.107 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ Q-8750 8.265 ms 0 - 9 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8295P 14.927 ms 48 - 53 MB NPU
decoder VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.339 ms 45 - 54 MB NPU
decoder VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile 8.307 ms 40 - 53 MB NPU
decoder VOICE_AI float Snapdragon® X2 Elite 6.846 ms 60 - 60 MB NPU
decoder VOICE_AI float Snapdragon® X Elite 10.576 ms 60 - 60 MB NPU
decoder VOICE_AI float Snapdragon® 8 Gen 3 Mobile 9.715 ms 37 - 45 MB NPU
decoder VOICE_AI float Snapdragon® 8 Gen 1 Mobile 14.144 ms 60 - 75 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 14.628 ms 60 - 130 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.274 ms 60 - 62 MB NPU
decoder VOICE_AI float Qualcomm® SA8775P 13.914 ms 44 - 53 MB NPU
decoder VOICE_AI float Qualcomm® SA8650P 13.914 ms 44 - 53 MB NPU
decoder VOICE_AI float Qualcomm® SA8255P 13.914 ms 44 - 53 MB NPU
decoder VOICE_AI float Qualcomm® QCS8450 14.144 ms 60 - 75 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 13.572 ms 60 - 129 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-X7181 10.576 ms 60 - 60 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ Q-8750 8.307 ms 40 - 53 MB NPU
decoder VOICE_AI float Qualcomm® SA8295P 15.108 ms 58 - 63 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 51.113 ms 129 - 142 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile 65.615 ms 82 - 89 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite 52.913 ms 132 - 132 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® X Elite 117.093 ms 254 - 254 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile 86.96 ms 128 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile 176.801 ms 125 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 147.391 ms 125 - 129 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) 116.576 ms 125 - 127 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® QCS8450 176.801 ms 125 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 141.52 ms 127 - 131 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-X7181 117.093 ms 254 - 254 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ Q-8750 65.615 ms 82 - 89 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 50.763 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile 65.801 ms 29 - 42 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® X2 Elite 52.624 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® X Elite 118.282 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile 85.303 ms 1 - 8 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile 172.287 ms 0 - 10 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 146.361 ms 0 - 57 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) 114.709 ms 0 - 4 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8775P 140.219 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8650P 140.219 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8255P 140.219 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® QCS8450 172.287 ms 0 - 10 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 140.417 ms 0 - 56 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-X7181 118.282 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ Q-8750 65.801 ms 29 - 42 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8295P 175.288 ms 0 - 5 MB NPU
encoder VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 50.886 ms 1 - 10 MB NPU
encoder VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile 65.633 ms 1 - 13 MB NPU
encoder VOICE_AI float Snapdragon® X2 Elite 53.015 ms 0 - 0 MB NPU
encoder VOICE_AI float Snapdragon® X Elite 118.324 ms 0 - 0 MB NPU
encoder VOICE_AI float Snapdragon® 8 Gen 3 Mobile 85.696 ms 1 - 8 MB NPU
encoder VOICE_AI float Snapdragon® 8 Gen 1 Mobile 171.546 ms 1 - 10 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 147.499 ms 0 - 57 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) 116.177 ms 1 - 4 MB NPU
encoder VOICE_AI float Qualcomm® SA8775P 139.985 ms 0 - 9 MB NPU
encoder VOICE_AI float Qualcomm® SA8650P 139.985 ms 0 - 9 MB NPU
encoder VOICE_AI float Qualcomm® SA8255P 139.985 ms 0 - 9 MB NPU
encoder VOICE_AI float Qualcomm® QCS8450 171.546 ms 1 - 10 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 141.699 ms 2 - 58 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-X7181 118.324 ms 0 - 0 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ Q-8750 65.633 ms 1 - 13 MB NPU
encoder VOICE_AI float Qualcomm® SA8295P 176.027 ms 1 - 6 MB NPU

License

  • The license for the original implementation of Whisper-Small can be found here.

References

Community

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support