Instructions to use litert-community/DeepSeek-R1-Distill-Qwen-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/DeepSeek-R1-Distill-Qwen-7B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/DeepSeek-R1-Distill-Qwen-7B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/DeepSeek-R1-Distill-Qwen-7B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).
Measured on device (edge-compat): Galaxy S26 · LiteRT-LM 0.16.0 · GPU · decode 8.4 tok/s · prefill 108 tok/s · TTFT 1.95 s · all 1243 ops delegated (2026-08-24); Raspberry Pi 5 · LiteRT-LM 0.16.1 · CPU, 4 threads · decode 1.8 tok/s · prefill 11 tok/s · TTFT 28.25 s (2026-09-01). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/deepseek-r1-distill-qwen-7b-q4-block32-ekv4096/CARD.md
DeepSeek-R1-Distill-Qwen-7B — LiteRT-LM (blockwise int4)
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
converted to the LiteRT-LM (.litertlm) format for on-device inference with
Google's LiteRT-LM runtime (the
engine behind the official litert-community/* models).
A reasoning model: it emits a <think> … </think> chain before the answer.
MIT-licensed (distilled onto an Apache-2.0 Qwen2.5 base). Converted with the
official upstream litert-torch — no fork, no custom code.
| File | DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm (~4.2 GB) |
| Quantization | int4 weights — blockwise (block 32) + OCTAV optimal-clipping, symmetric; embedding INT8 |
| Compute | integer |
| Context (KV cache) | 4096 |
| Base model | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B |
| Platforms | Desktop (Mac) ✓ · high-RAM (12 GB+) Android ✓ · iPhone / 8 GB phones ✗ (4 GB exceeds the budget) |
2026-10-06: GPU graph rewrite. DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm was replaced by a file with a rewritten prefill/decode graph. In each layer, the two DYNAMIC_UPDATE_SLICE KV-cache writes became one odml.cache_update composite, and the BATCH_MATMUL attention products became odml.runtime_bmm composites. This is the graph shape litert-torch exports with --apply_gpu_composites; the previous file was exported without that flag. The graph also gains a signature input, param_tensor (INT32[1,1,1,7]), whose start/end values the runtime fills. The attention mask stays a FLOAT32 input. A 16-token prefill signature, prefill_16, was added as a copy of prefill_128 that shares its weight buffers. Every other bundle section, and the weight region of the graph, is byte-identical to the previous file (checked per section and per buffer). The new file is 4,533,682,160 bytes, sha256 050787a59d80acf064ed71209e45a28a7ae734a06f409f271751f254643216e1. The previous file (Hub revision 96991f3653ce) was 4,531,978,224 bytes, sha256 511d59c11704f7ab39b9cb3a0eef1a88f8c05673d0b74bcb4b19c15845ab4a7f. On CPU, prefill_16 gives output bit-identical to prefill_128 on the same tokens. With and without prefill_16, the answers to 9 gate prompts (an 8-question sanity gate plus one long prompt) are byte-identical, on CPU and on the Mac GPU with fp32 activations. On the Mac GPU with fp32 activations, the rewritten graph without prefill_16 also matches the previous file byte for byte on all 8 questions. On the Mac GPU with the default fp16 activations, 21 of 30 GSM8K answer texts stay byte-identical. Both files answer 8/8 sanity-gate questions. On 30 GSM8K questions (max tokens 2048, greedy) the new file scores 28/30 against 27/30 for the previous file; the one-question difference has McNemar p = 1.0. The Raspberry Pi 5 rows further down were measured on the previous file. The commands are under Conversion.
Mac Studio M4 Max, litert-lm 0.17.1 benchmark --cache no, both files measured in the same window (median of 3 processes on GPU, 2 on CPU; GPU is WebGPU on Metal with fp16 activations, CPU is XNNPACK):
| Measure | Previous file | New file | New / previous |
|---|---|---|---|
| GPU decode (256-token prompt, 256 decode tokens) | 65.0 tok/s | 79.1 tok/s | ×1.217 |
| GPU prefill (256-token prompt) | 644 tok/s | 675 tok/s | ×1.049 |
| GPU TTFT (16-token prompt) | 0.220 s | 0.046 s | ×0.210 |
| CPU decode (256-token prompt, 256 decode tokens) | 18.1 tok/s | 18.2 tok/s | ×1.005 |
| CPU TTFT (16-token prompt) | 2.893 s | 1.192 s | ×0.412 |
Usage
litert_lm_main --model_path DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm --backend gpu \
--input_prompt "If a train travels 60 km in 45 minutes, what is its speed in km/h?"
The .litertlm bundle carries the tokenizer and the DeepSeek prompt template
(<|User|> / <|Assistant|>, stop token <|end▁of▁sentence|>). The assistant
opens a <think> block, reasons step by step, then gives the final answer
(commonly in \boxed{}).
Performance
All rows are for the current file (the 2026-10-06 file).
Mac Studio M4 Max. litert-lm 0.17.1 (pip) CLI, litert-lm benchmark --cache no, 256 prompt / 256 decode tokens. Each value is the median of separate processes: 3 on GPU, 2 on CPU. The 16-token TTFT comes from separate runs with a 16-token prompt and 32 decode tokens.
| Backend | Prefill (256 tokens) | Decode (256 tokens) | TTFT (256-token prompt) | TTFT (16-token prompt) | Init |
|---|---|---|---|---|---|
| GPU (WebGPU on Metal, fp16 activations) | 675 tok/s | 79.1 tok/s | 0.39 s | 0.046 s | 5.0 s |
| CPU (XNNPACK) | 65 tok/s | 18.2 tok/s | 4.84 s | 1.192 s | 6.0 s |
Samsung Galaxy S26 (SM-S942Q, Android 16). litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18, 1024 prompt / 256 decode tokens. The GPU runs use --disable_cache=true (cold init); the CPU runs use the XNNPACK weight cache, which init writes. Each run is one process: 3 iterations on GPU, 2 on CPU, and 3 in each separate 16-token run (16 prompt / 32 decode tokens). Each cell is the range over those iterations. In the 1024-token GPU run the first iteration is the fastest and the third the slowest (cold to warm). In the 16-token runs the first iteration has the longest TTFT. Init and peak memory (the process high-water mark, VmHWM) are from the 1024-token run.
| Backend | Prefill (1024 tokens) | Decode (256 tokens) | TTFT (1024-token prompt) | TTFT (16-token prompt) | Init | Peak |
|---|---|---|---|---|---|---|
| GPU (OpenCL, fp16 activations) | 87–186 tok/s | 8.5–10.3 tok/s | 5.61–11.89 s | 0.21–0.23 s | 27.8 s | 2.00 GB |
| CPU (XNNPACK, 4 threads) | 31–36 tok/s | 8.2–8.6 tok/s | 28.65–33.33 s | 0.43–2.72 s | 13.9 s | 5.66 GB |
At 4.5 GB this build does not fit an 8 GB phone — its weights section exceeds the iOS mmap budget — so there is no iPhone row. It does load and generate on a 12 GB Android phone: see the Galaxy S26 section below.
Accuracy note
GSM8K (n=100, greedy, 0-shot, identical prompt + answer-extraction; max_new_tokens=2048
to fit the reasoning chain).
| Configuration | GSM8K |
|---|---|
| bf16 (reference) | 88.0% |
| This model — LiteRT int4 (BOCTAV4) | 87.0% |
LiteRT int4 is at parity — −1.0 pt vs bf16. The reasoning behavior is fully preserved through 4-bit quantization; the shallow-wide Qwen2 (28 layers) absorbs int4 rounding cleanly.
Galaxy S26 — GPU backend
The published bundle runs on the Android GPU backend and generates.
On a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18, the three transformer signatures run fully on the OpenCL GPU delegate (LITERT_CL), one partition each:
| signature | nodes on LITERT_CL |
|---|---|
prefill_128 |
1268 / 1268 |
prefill_16 |
1268 / 1268 |
decode |
1159 / 1159 |
XNNPACK takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128. On the GPU the file answered 4 of 5 gate prompts correctly. Speed and peak memory are in the Galaxy S26 table under Performance.
GPU wiring, including the Gallery import toggle: GPU guide.
Conversion
Converted with the official upstream litert-torch
export_hf (clean git worktree at upstream/main, dev-fork patches excluded).
Qwen2ForCausalLM rides the stock converter with no custom code. int4 recipe =
blockwise (block 32) + OCTAV with INT8 embedding (externalized into its own
bundle section); KV cache 4096.
Graph rewrite (2026-10-06). The current file was built from the previous file (Hub revision 96991f3653ce) with two tools in tools/gpu_graph/ of hf-to-litertlm:
python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
python tools/gpu_graph/prefill_bucket_clone.py build retrofit.litertlm new.litertlm --lengths 16 --source prefill_128 --report bucket.json
previous.litertlm is the previous file, new.litertlm the result published here as DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm; LITERT_LM_CLI names the litert-lm CLI the scripts use to unpack and pack the bundle.
gpu_graph_retrofit.py changes only the prefill/decode TFLite graph, and every original weight buffer keeps its bytes. With --mask add_bcast, the attention mask is applied as an ADD on a [bk, g, T, C] view of the logits, with the FLOAT32 mask broadcast. prefill_bucket_clone.py copies an existing prefill signature (prefill_128) at a new length (16); the copy shares every weight buffer, so the file grows by graph structure only. A rebuild with these tools has every bundle section byte-identical to the shipped file. Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.
Training data & PII
This is a weights-exact format conversion of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B; no new training was performed. It is a Qwen2.5-7B-family base supervised-fine-tuned by DeepSeek on ~800K reasoning traces generated by DeepSeek-R1; the distillation set is model-generated and the Qwen base pretraining corpus is web-derived and not fully disclosed. Web-derived data may incidentally contain PII; none was deliberately collected and this format conversion adds none. Apply your own content/PII filtering before deployment. See the base model card for details.
2026-08-31 — thought channel declared (metadata only, weights unchanged)
DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm now declares the reasoning channel in its metadata (LlmMetadata.channels: channel name thought, markers <think>…</think> exactly as this model emits them). Without the declaration the runtime has no way to tell the reasoning apart from the answer: the raw thinking streamed inline into the visible text, and a thinking_token_budget was silently ignored (the API returns OK and only logs a warning). With the channel declared, LiteRT-LM returns the reasoning separated in channels["thought"] and the thinking budget takes effect.
Metadata-only change: every section of the bundle except LlmMetadata is byte-identical to the previous file (verified per section, tokenizer included), so the weights, the graph, the tokenizer and the chat template are unchanged and the speed and accuracy numbers on this card still describe this file — only the file's own sha256 differs. Verified on the LiteRT-LM runtime (litert-lm-api 0.16.1): the visible answer stays clean, the reasoning lands in channels["thought"], and on one bundle of this batch thinking_token_budget=16 was confirmed to truncate the reasoning at exactly 16 tokens where it was a no-op before. If you downloaded before 2026-08-31, re-download to get the channel-aware file.
Raspberry Pi 5 (CPU) (previous file)
This row was measured on the previous file (sha256 511d59c11704f7ab39b9cb3a0eef1a88f8c05673d0b74bcb4b19c15845ab4a7f) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical in the new file, and the new file's Mac CPU speed is ×1.005 of the previous file's in both decode and prefill, so the row is expected to hold.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (min–max in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
DeepSeek-R1-Distill-Qwen-7B_q4_block32_ekv4096.litertlm |
11.0 (10.4–11.0) | 1.8 (1.7–1.8) | 28.3 s | 5.4 GB |
License
MIT (model weights), inherited from deepseek-ai/DeepSeek-R1-Distill-Qwen-7B; the Qwen2.5 base is Apache-2.0. Commercial use and derivatives permitted.
- Downloads last month
- 845
Model tree for litert-community/DeepSeek-R1-Distill-Qwen-7B
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B