TokForge
- Website: https://tokforge.ai
- Discord: https://discord.gg/Acv3CBtfVm
- Google Play: https://play.google.com/store/apps/details?id=dev.tokforge
- iOS TestFlight: https://testflight.apple.com/join/jnufjzRr
Runs on-device in the TokForge app.
TokForge LLM QNN NPU packs
NPU packs for the Faster long prompts (NPU) setting in TokForge on Snapdragon phones. The phone's NPU reads your long prompt; the CPU writes the reply with the same model as before.
Original models by the Qwen team. The Qwen3 4B pack is built from huihui-ai/Huihui-Qwen3-4B-abliterated-v2, an abliterated variant of Qwen3 4B. These packs only change the runtime format for the phone's NPU.
What you get
- Faster first replies on long prompts: up to about 3x faster, depending on the model and chip (table below).
- Replies still come from your model: the CPU writes every reply with the model you downloaded; the pack only reads the prompt.
- Checked before use: every file is verified by sha256 on download, and TokForge runs a quick on-device check before it uses the NPU. If the check does not pass, prompts are read on the CPU as before.
- Download only what your phone needs: one folder per model and chip, fetched only when you turn the setting on.
Supported phones
| Chip | NPU | Folder |
|---|---|---|
| Snapdragon 8 Gen 3 (SM8650) | Hexagon V75 | v75/ |
| Snapdragon 8 Elite Gen 5 (SM8850) | Hexagon V81 | v81/ |
Packs
| Folder | Model in TokForge | Works with this exact download | Size per chip | First reply on a long prompt |
|---|---|---|---|---|
Qwen3-1.7B-MNN-HQQ-catalogpair |
Qwen3 1.7B | darkmaniac7/Qwen3-1.7B-MNN-HQQ @ 15ab6d1d |
1.75 GB | up to about 2x faster |
Qwen3-1.7B-MNN-catalogpair |
Qwen3 1.7B (earlier download) | taobao-mnn/Qwen3-1.7B-MNN @ 1b12fe7d |
1.75 GB | up to about 3x faster |
Qwen3-4B-abliterated-MNN-catalogpair |
Qwen3 4B abliterated | darkmaniac7/Qwen3-4B-abliterated-MNN @ edba0fdb |
4.09 GB | up to about 3x faster |
Qwen3-0.6B-MNN-catalogpair |
Qwen3 0.6B (worker model) | taobao-mnn/Qwen3-0.6B-MNN @ 34dfccda |
0.62 GB | up to about 1.4x faster |
A pack is matched to its model download by file hashes. With any other download it is not used.
Files
<pack>/MANIFEST.json every file per chip with sha256 and size, and the exact download it pairs with
<pack>/<chip>/config_qnn.json runtime config for the NPU session
<pack>/<chip>/qnn/graph*.bin Qualcomm QNN context binaries (the model body, compiled for that NPU)
<pack>/<chip>/qnn/llm.mnn MNN graph that drives the binaries
<pack>/<chip>/qnn/llm_config.json
LICENSE ATTRIBUTION.txt MODIFICATIONS.txt
Usage with TokForge
- Install TokForge 1.3.8 or later on a supported phone.
- Download one of the models above from the app's model list.
- Open the model's settings and turn on Faster long prompts (NPU). TokForge downloads the pack for your chip and asks for a restart before first use.
The NPU runtime libraries come from
darkmaniac7/TokForge-SD15-QNN-NPU (qnn_runtime/2.40).
How the packs are made
Each pack is exported from its paired download's own weights: MNN's loader dequantizes them, and MNN llmexport writes
8-bit per-channel weights for the NPU graph. Rotary position tables are fed from fp32 tables, as on the CPU. The body is
compiled offline into QNN context binaries with Qualcomm AI Runtime 2.40. Source training weights were not edited. See
MODIFICATIONS.txt.
Limitations
- Snapdragon 8 Gen 3 and 8 Elite Gen 5 only. On other chips TokForge reads prompts on the CPU.
- Long prompts only: the NPU reads prompts of about 512 to 2,000 tokens; shorter prompts stay on the CPU.
- Runtime files for TokForge and MNN, not a Transformers checkpoint.
- The models are small: they can still get facts wrong. Check important answers.
License and attribution
Apache License 2.0 (LICENSE). The original models are by the Qwen team; the Qwen3 4B pack builds on huihui-ai's
abliterated variant. All source authors keep their copyrights; no endorsement is claimed. Details in
ATTRIBUTION.txt and MODIFICATIONS.txt.
Community
- Website: tokforge.ai
- Discord: Join our Discord