TokForge

Runs on-device in the TokForge app.

TokForge Acceleration Pack โ€” Qwen3.5 Draft (Deprecated)

Deprecated: this Qwen3.5-0.8B draft bundle is preserved for reproducibility, but the newer Qwen3-0.6B TokForge draft line is the practical default.

Why it is deprecated

Two reasons:

  • Cross-family draft pairing does not attach: Qwen3.5 and Qwen3 use different vocabularies (248,320 vs 151,936), so this 0.8B draft cannot pair with Qwen3 targets, and the Qwen3-0.6B drafts cannot pair with Qwen3.5 targets.
  • In our testing, Qwen3.5 targets did not get reliable speedups from speculative decoding. The same-family Qwen3-0.6B draft line with dense Qwen3 targets measured +34% to +43% faster decode in chat workloads on our test devices.

Use instead

Performance Notes

This bundle is preserved because it was an important research step, but it is not the current practical winner. The measured wins live in the same-family Qwen3-0.6B draft line with dense Qwen3 targets; this Qwen3.5 draft lane did not show reliable speedups in our testing.

What is included

  • llm.mnn
  • llm.mnn.weight
  • llm_config.json
  • tokenizer file(s)

Usage

This repo is for TokForge / MNN users who specifically want to reproduce the older Qwen3.5-0.8B draft path.

Limitations and Intended Use

  • Deprecated for normal use.
  • Cross-family drafting is not possible (vocabulary mismatch), and same-family Qwen3.5 pairing did not show reliable speedups in our testing.
  • Keep using this only if you need reproducibility for old experiments.
Downloads last month
138
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for darkmaniac7/TokForge-AccelerationPack-Qwen35-Draft

Finetuned
(307)
this model