TokForge
- Website: https://tokforge.ai
- Discord: https://discord.gg/Acv3CBtfVm
- Google Play: https://play.google.com/store/apps/details?id=dev.tokforge
- iOS TestFlight: https://testflight.apple.com/join/jnufjzRr
Runs on-device in the TokForge app.
TokForge Acceleration Pack โ Qwen3.5 Draft (Deprecated)
Deprecated: this
Qwen3.5-0.8Bdraft bundle is preserved for reproducibility, but the newerQwen3-0.6BTokForge draft line is the practical default.
Why it is deprecated
Two reasons:
- Cross-family draft pairing does not attach: Qwen3.5 and Qwen3 use different vocabularies (248,320 vs 151,936), so this 0.8B draft cannot pair with Qwen3 targets, and the Qwen3-0.6B drafts cannot pair with Qwen3.5 targets.
- In our testing, Qwen3.5 targets did not get reliable speedups from speculative decoding. The same-family
Qwen3-0.6Bdraft line with dense Qwen3 targets measured +34% to +43% faster decode in chat workloads on our test devices.
Use instead
- Recommended replacement: TokForge-AccelerationPack-Draft
- Collection: TokForge Mobile Draft Models
Performance Notes
This bundle is preserved because it was an important research step, but it is not the current practical winner. The measured wins live in the same-family Qwen3-0.6B draft line with dense Qwen3 targets; this Qwen3.5 draft lane did not show reliable speedups in our testing.
What is included
llm.mnnllm.mnn.weightllm_config.json- tokenizer file(s)
Usage
This repo is for TokForge / MNN users who specifically want to reproduce the older Qwen3.5-0.8B draft path.
Limitations and Intended Use
- Deprecated for normal use.
- Cross-family drafting is not possible (vocabulary mismatch), and same-family Qwen3.5 pairing did not show reliable speedups in our testing.
- Keep using this only if you need reproducibility for old experiments.
- Downloads last month
- 138