TokForge

Runs on-device in the TokForge app.

TokForge LLM QNN NPU packs

NPU packs for the Faster long prompts (NPU) setting in TokForge on Snapdragon phones. The phone's NPU reads your long prompt; the CPU writes the reply with the same model as before.

Original models by the Qwen team. The Qwen3 4B pack is built from huihui-ai/Huihui-Qwen3-4B-abliterated-v2, an abliterated variant of Qwen3 4B. These packs only change the runtime format for the phone's NPU.

What you get

  • Faster first replies on long prompts: up to about 3x faster, depending on the model and chip (table below).
  • Replies still come from your model: the CPU writes every reply with the model you downloaded; the pack only reads the prompt.
  • Checked before use: every file is verified by sha256 on download, and TokForge runs a quick on-device check before it uses the NPU. If the check does not pass, prompts are read on the CPU as before.
  • Download only what your phone needs: one folder per model and chip, fetched only when you turn the setting on.

Supported phones

Chip NPU Folder
Snapdragon 8 Gen 3 (SM8650) Hexagon V75 v75/
Snapdragon 8 Elite Gen 5 (SM8850) Hexagon V81 v81/

Packs

Folder Model in TokForge Works with this exact download Size per chip First reply on a long prompt
Qwen3-1.7B-MNN-HQQ-catalogpair Qwen3 1.7B darkmaniac7/Qwen3-1.7B-MNN-HQQ @ 15ab6d1d 1.75 GB up to about 2x faster
Qwen3-1.7B-MNN-catalogpair Qwen3 1.7B (earlier download) taobao-mnn/Qwen3-1.7B-MNN @ 1b12fe7d 1.75 GB up to about 3x faster
Qwen3-4B-abliterated-MNN-catalogpair Qwen3 4B abliterated darkmaniac7/Qwen3-4B-abliterated-MNN @ edba0fdb 4.09 GB up to about 3x faster
Qwen3-0.6B-MNN-catalogpair Qwen3 0.6B (worker model) taobao-mnn/Qwen3-0.6B-MNN @ 34dfccda 0.62 GB up to about 1.4x faster

A pack is matched to its model download by file hashes. With any other download it is not used.

Files

<pack>/MANIFEST.json            every file per chip with sha256 and size, and the exact download it pairs with
<pack>/<chip>/config_qnn.json   runtime config for the NPU session
<pack>/<chip>/qnn/graph*.bin    Qualcomm QNN context binaries (the model body, compiled for that NPU)
<pack>/<chip>/qnn/llm.mnn       MNN graph that drives the binaries
<pack>/<chip>/qnn/llm_config.json
LICENSE  ATTRIBUTION.txt  MODIFICATIONS.txt

Usage with TokForge

  1. Install TokForge 1.3.8 or later on a supported phone.
  2. Download one of the models above from the app's model list.
  3. Open the model's settings and turn on Faster long prompts (NPU). TokForge downloads the pack for your chip and asks for a restart before first use.

The NPU runtime libraries come from darkmaniac7/TokForge-SD15-QNN-NPU (qnn_runtime/2.40).

How the packs are made

Each pack is exported from its paired download's own weights: MNN's loader dequantizes them, and MNN llmexport writes 8-bit per-channel weights for the NPU graph. Rotary position tables are fed from fp32 tables, as on the CPU. The body is compiled offline into QNN context binaries with Qualcomm AI Runtime 2.40. Source training weights were not edited. See MODIFICATIONS.txt.

Limitations

  • Snapdragon 8 Gen 3 and 8 Elite Gen 5 only. On other chips TokForge reads prompts on the CPU.
  • Long prompts only: the NPU reads prompts of about 512 to 2,000 tokens; shorter prompts stay on the CPU.
  • Runtime files for TokForge and MNN, not a Transformers checkpoint.
  • The models are small: they can still get facts wrong. Check important answers.

License and attribution

Apache License 2.0 (LICENSE). The original models are by the Qwen team; the Qwen3 4B pack builds on huihui-ai's abliterated variant. All source authors keep their copyrights; no endorsement is claimed. Details in ATTRIBUTION.txt and MODIFICATIONS.txt.

Community

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darkmaniac7/TokForge-LLM-QNN-NPU

Finetuned
Qwen/Qwen3-0.6B
Quantized
(477)
this model