ScamScan Android NPU runtime assets
Runtime model package for ScamScan code commit 37af8ac. Parakeet TDT 0.6B v3 is
the default speech model and Whisper Small is the alternative. Both run on the
Snapdragon NPU and feed the existing fine-tuned MiniLM scam classifier. If an NPU
session fails to load, the app falls back to Whisper on the CPU with the same
classifier and shows the reason. Bodhan is not part of this package. The
repository is gated: request access on this page, then use a token that can read
public gated repositories.
Required files and hardware
models/parakeet_qnn/: Parakeet encoder/decoder ONNX wrappers and HTP contexts.models/whisper_small_qnn/: Whisper Small wrappers, HTP contexts and metadata.android/app/src/main/assets/: shared classifier, frontend, vocabularies, Whisper generation settings and keyword rules.runtime-manifest.json: exact sizes and SHA-256 checksums for every runtime file.models/minilm_classifier/: the trained PyTorch MiniLM (BertForSequenceClassification, label 0 benign, label 1 scam) thatscamscan_classifier_htp.onnxwas exported from, with its config and tokenizer. The app does not download it; use it to re-export withpipeline/export_classifier_htp.pyor to fine-tune further.
scam_keywords_in.json, parakeet_vocab.txt and whisper_small_generation.json
are also tracked in the code repository. The keyword rules here match code commit
37af8ac. The two text files here use CRLF line endings; their content matches the
code repository, and the app reads either.
The CPU fallback uses the root-level whisper_encoder.onnx and
whisper_decoder.onnx from revision 37464c16fa9e276a9c72d1c0374e6b76a7518ff4.
The root-level scamscan_classifier.onnx is a legacy graph that returns the same
logits for every input. The app rejects it at load time; the NPU pipelines and the
CPU fallback both use android/app/src/main/assets/scamscan_classifier_htp.onnx.
The other root-level files remain for older code.
Validated on Qualcomm SM8850, Hexagon HTP v81, QAIRT 2.50.0 and ONNX Runtime Android QNN 1.29.0. These precompiled contexts are device-specific; compatibility with other chips is not established. Both ASR neural models and the classifier use strict QNN sessions. Signal processing, tokenization and decoding control run on CPU; Parakeet's small STFT/mel ONNX frontend also runs on CPU.
Restore into the code checkout
From the ScamScan source repository root:
python3 scripts/setup_model_assets.py
The script downloads runtime-manifest.json at a pinned revision, restores every
listed file that Git does not track, checks each file's size and SHA-256, and
fetches the two CPU fallback Whisper files. --skip-npu leaves out the 1.8 GB
context packages.
To download by hand, pin a revision and exclude the files the code repository tracks so the download does not overwrite them:
hf download Fugi1128/ScamScan --revision <Hub commit SHA> --include "models/parakeet_qnn/*" --include "models/whisper_small_qnn/*" --include "android/app/src/main/assets/*" --include "runtime-manifest.json" --exclude "android/app/src/main/assets/scam_keywords_in.json" --exclude "android/app/src/main/assets/parakeet_vocab.txt" --exclude "android/app/src/main/assets/whisper_small_generation.json" --local-dir .
Then verify each file's size and SHA-256 against runtime-manifest.json.
Build the Android APK with the restored assets and install it without
uninstalling the previous app. Then provision the contexts with the source
repository scripts (use python on Windows):
python3 pipeline/provision_parakeet_qnn.py
python3 pipeline/provision_whisper_qnn.py
They need Python 3.9 or newer and a debuggable installed app. They find adb
through $ADB, PATH, ANDROID_HOME or ANDROID_SDK_ROOT; pass --adb <path>
to use another one. Each file is verified on the device. The APK alone does not
provision the large ASR context binaries. Qualcomm SDK and runtime libraries are
not distributed in this repository; obtain the required runtime/toolchain
separately as configured in the app build.
Validation and limitations
With code commit 37af8ac, end-to-end calls from the ScamScan caller portal to the
phone ran Parakeet and the classifier on the NPU at 83 to 106 ms per audio chunk.
Only the classifier moves the risk score. A recorded digital-arrest scam call
reached 97% (Critical) within 14 seconds, and a recorded food-delivery call stayed
at 0% under a police caller name.
The classifier in this revision was retrained on public call datasets and generated Indian-context scam calls. On held-out calls it missed 2.0% of scam windows (the previous classifier missed 46.6%) and flagged 1.4% of benign windows (previously 2.8%). With the app's smoothing, every held-out scam call raised an alert and 313 of 314 reached Critical, while 2.7% of benign calls reached a false Critical (previously 7.7%).
Foreground device tests also passed three Parakeet/Whisper round trips with ASR and classification after each switch. Tests used English recordings and synthetic fixtures; this is not an independent human-call WER benchmark or a sustained background test. The classifier is a development scam-detection model, not a guarantee of safety.
Attribution and licenses
This repository is a bundle; licenses apply per component rather than a single blanket license for all files.
- Parakeet: NVIDIA parakeet-tdt-0.6b-v3, CC BY 4.0. Frontend, vocabulary and source ONNX export from istupakov/parakeet-tdt-0.6b-v3-onnx. Changes: fixed-shape specialization, FP16 conversion and QAIRT HTP compilation.
- Whisper: OpenAI Whisper, MIT license.
Compiled package from Qualcomm AI Hub Models whisper_small release v0.63.0,
whisper_small-precompiled_qnn_onnx-float-qualcomm_snapdragon_8_elite_gen5_for_galaxy.zip. Original archive SHA-256:c02d8e86b541f5b259b3b0f2b300b400a6a2e2828ceff10b668f3b3909e9f074. - MiniLM classifier: fine-tuned from sentence-transformers/all-MiniLM-L6-v2 (Apache 2.0) and exported to a static QNN-compatible ONNX graph. Training data: the BothBosu scam, single-agent, multi-agent and YouTube conversation sets (Apache 2.0), the scam half of shakeleoatmeal/phone-scam-detection-synthetic (MIT), Ngadou/social-engineering-convo (Apache 2.0), the Talkmap telecom and banking corpora (MIT), Lakshan2003/customer-support-client-agent-conversations (MIT), the AppTek call-centre dialogues (CC BY-SA 4.0), Google Taskmaster (CC BY 4.0), 3nesdeniz/english-daily-dialogues-10k (CC BY 4.0), banking77 (MIT) and ScamScan's generated Indian-context examples.