|
Download README.md from Fugi1128/ScamScan: direct link, hf CLI and curl.
- Browser
- Download file 7.31 kB
-
https://huggingface.co/Fugi1128/ScamScan/resolve/main/README.md
- Command line
-
hf download hf://Fugi1128/ScamScan/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/Fugi1128/ScamScan/resolve/main/README.md
7.31 kB
| language: | |
| - en | |
| tags: | |
| - onnx | |
| - android | |
| - qnn | |
| - automatic-speech-recognition | |
| - text-classification | |
| license: other | |
| license_name: component-specific | |
| license_link: https://huggingface.co/Fugi1128/ScamScan/blob/main/README.md#attribution-and-licenses | |
| # ScamScan Android NPU runtime assets | |
| Runtime model package for ScamScan code commit `37af8ac`. Parakeet TDT 0.6B v3 is | |
| the default speech model and Whisper Small is the alternative. Both run on the | |
| Snapdragon NPU and feed the existing fine-tuned MiniLM scam classifier. If an NPU | |
| session fails to load, the app falls back to Whisper on the CPU with the same | |
| classifier and shows the reason. Bodhan is not part of this package. The | |
| repository is gated: request access on this page, then use a token that can read | |
| public gated repositories. | |
| ## Required files and hardware | |
| - `models/parakeet_qnn/`: Parakeet encoder/decoder ONNX wrappers and HTP contexts. | |
| - `models/whisper_small_qnn/`: Whisper Small wrappers, HTP contexts and metadata. | |
| - `android/app/src/main/assets/`: shared classifier, frontend, vocabularies, | |
| Whisper generation settings and keyword rules. | |
| - `runtime-manifest.json`: exact sizes and SHA-256 checksums for every runtime file. | |
| - `models/minilm_classifier/`: the trained PyTorch MiniLM (`BertForSequenceClassification`, | |
| label 0 benign, label 1 scam) that `scamscan_classifier_htp.onnx` was exported from, | |
| with its config and tokenizer. The app does not download it; use it to re-export | |
| with `pipeline/export_classifier_htp.py` or to fine-tune further. | |
| `scam_keywords_in.json`, `parakeet_vocab.txt` and `whisper_small_generation.json` | |
| are also tracked in the code repository. The keyword rules here match code commit | |
| `37af8ac`. The two text files here use CRLF line endings; their content matches the | |
| code repository, and the app reads either. | |
| The CPU fallback uses the root-level `whisper_encoder.onnx` and | |
| `whisper_decoder.onnx` from revision `37464c16fa9e276a9c72d1c0374e6b76a7518ff4`. | |
| The root-level `scamscan_classifier.onnx` is a legacy graph that returns the same | |
| logits for every input. The app rejects it at load time; the NPU pipelines and the | |
| CPU fallback both use `android/app/src/main/assets/scamscan_classifier_htp.onnx`. | |
| The other root-level files remain for older code. | |
| Validated on Qualcomm SM8850, Hexagon HTP v81, QAIRT 2.50.0 and ONNX Runtime | |
| Android QNN 1.29.0. These precompiled contexts are device-specific; compatibility | |
| with other chips is not established. Both ASR neural models and the classifier | |
| use strict QNN sessions. Signal processing, tokenization and decoding control | |
| run on CPU; Parakeet's small STFT/mel ONNX frontend also runs on CPU. | |
| ## Restore into the code checkout | |
| From the ScamScan source repository root: | |
| ```sh | |
| python3 scripts/setup_model_assets.py | |
| ``` | |
| The script downloads `runtime-manifest.json` at a pinned revision, restores every | |
| listed file that Git does not track, checks each file's size and SHA-256, and | |
| fetches the two CPU fallback Whisper files. `--skip-npu` leaves out the 1.8 GB | |
| context packages. | |
| To download by hand, pin a revision and exclude the files the code repository | |
| tracks so the download does not overwrite them: | |
| ```sh | |
| hf download Fugi1128/ScamScan --revision <Hub commit SHA> --include "models/parakeet_qnn/*" --include "models/whisper_small_qnn/*" --include "android/app/src/main/assets/*" --include "runtime-manifest.json" --exclude "android/app/src/main/assets/scam_keywords_in.json" --exclude "android/app/src/main/assets/parakeet_vocab.txt" --exclude "android/app/src/main/assets/whisper_small_generation.json" --local-dir . | |
| ``` | |
| Then verify each file's size and SHA-256 against `runtime-manifest.json`. | |
| Build the Android APK with the restored assets and install it without | |
| uninstalling the previous app. Then provision the contexts with the source | |
| repository scripts (use `python` on Windows): | |
| ```sh | |
| python3 pipeline/provision_parakeet_qnn.py | |
| python3 pipeline/provision_whisper_qnn.py | |
| ``` | |
| They need Python 3.9 or newer and a debuggable installed app. They find `adb` | |
| through `$ADB`, `PATH`, `ANDROID_HOME` or `ANDROID_SDK_ROOT`; pass `--adb <path>` | |
| to use another one. Each file is verified on the device. The APK alone does not | |
| provision the large ASR context binaries. Qualcomm SDK and runtime libraries are | |
| not distributed in this repository; obtain the required runtime/toolchain | |
| separately as configured in the app build. | |
| ## Validation and limitations | |
| With code commit `37af8ac`, end-to-end calls from the ScamScan caller portal to the | |
| phone ran Parakeet and the classifier on the NPU at 83 to 106 ms per audio chunk. | |
| Only the classifier moves the risk score. A recorded digital-arrest scam call | |
| reached 97% (Critical) within 14 seconds, and a recorded food-delivery call stayed | |
| at 0% under a police caller name. | |
| The classifier in this revision was retrained on public call datasets and | |
| generated Indian-context scam calls. On held-out calls it missed 2.0% of scam | |
| windows (the previous classifier missed 46.6%) and flagged 1.4% of benign windows | |
| (previously 2.8%). With the app's smoothing, every held-out scam call raised an | |
| alert and 313 of 314 reached Critical, while 2.7% of benign calls reached a false | |
| Critical (previously 7.7%). | |
| Foreground device tests also passed three Parakeet/Whisper round trips with ASR | |
| and classification after each switch. Tests used English recordings and synthetic | |
| fixtures; this is not an independent human-call WER benchmark or a sustained | |
| background test. The classifier is a development scam-detection model, not a | |
| guarantee of safety. | |
| ## Attribution and licenses | |
| This repository is a bundle; licenses apply per component rather than a single | |
| blanket license for all files. | |
| - Parakeet: NVIDIA [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), CC BY 4.0. | |
| Frontend, vocabulary and source ONNX export from | |
| [istupakov/parakeet-tdt-0.6b-v3-onnx](https://huggingface.co/istupakov/parakeet-tdt-0.6b-v3-onnx/tree/8f23f0c03c8761650bdb5b40aaf3e40d2c15f1ce). | |
| Changes: fixed-shape specialization, FP16 conversion and QAIRT HTP compilation. | |
| - Whisper: [OpenAI Whisper](https://github.com/openai/whisper), MIT license. | |
| Compiled package from Qualcomm AI Hub Models whisper_small release v0.63.0, | |
| `whisper_small-precompiled_qnn_onnx-float-qualcomm_snapdragon_8_elite_gen5_for_galaxy.zip`. | |
| Original archive SHA-256: `c02d8e86b541f5b259b3b0f2b300b400a6a2e2828ceff10b668f3b3909e9f074`. | |
| - MiniLM classifier: fine-tuned from | |
| [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | |
| (Apache 2.0) and exported to a static QNN-compatible ONNX graph. Training data: | |
| the BothBosu scam, single-agent, multi-agent and YouTube conversation sets | |
| (Apache 2.0), the scam half of shakeleoatmeal/phone-scam-detection-synthetic | |
| (MIT), Ngadou/social-engineering-convo (Apache 2.0), the Talkmap telecom and | |
| banking corpora (MIT), Lakshan2003/customer-support-client-agent-conversations | |
| (MIT), the AppTek call-centre dialogues (CC BY-SA 4.0), Google Taskmaster | |
| (CC BY 4.0), 3nesdeniz/english-daily-dialogues-10k (CC BY 4.0), banking77 (MIT) | |
| and ScamScan's generated Indian-context examples. | |