Instructions to use litert-community/VibeThinker-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/VibeThinker-3B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/VibeThinker-3B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/VibeThinker-3B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).
Measured on device (edge-compat): Galaxy S26 Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 13.9 tok/s Β· prefill 157 tok/s Β· TTFT 1.36 s Β· all 1603 ops delegated (2026-08-24); Pixel 8a Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 7.6 tok/s Β· prefill 13 tok/s Β· TTFT 1.48 s Β· all 1603 ops delegated (2026-08-17); Mac Studio M4 Max Β· LiteRT-LM 0.14.0 Β· GPU Β· decode 93.2 tok/s Β· prefill 1359 tok/s Β· TTFT 199 ms (2026-07-23). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/vibethinker-3b/CARD.md
VibeThinker-3B β LiteRT-LM (blockwise int4)
WeiboAI/VibeThinker-3B converted to the
LiteRT-LM (.litertlm) format for on-device inference with Google's
LiteRT-LM runtime (the engine behind the official
litert-community/* models).
VibeThinker-3B is a dense 3B math/reasoning model (Qwen2ForCausalLM, 36 layers) β it solves
problems with an inline chain-of-thought and is strong at arithmetic and math word problems. Standard
Qwen2 architecture, so it rides the existing converter and runtime directly.
| File | model.litertlm β int4 block 32 (~1.9 GB) |
| Quantization | int4 weights (symmetric) + OCTAV optimal-clipping; embeddings INT8 (externalized section) |
| Compute | integer |
| Context (KV cache) | 4096 |
| Base model | WeiboAI/VibeThinker-3B |
2026-10-06: GPU graph rewrite. Only the prefill/decode graph of model.litertlm changed. Its two DYNAMIC_UPDATE_SLICE KV-cache writes per layer are now one STABLEHLO_COMPOSITE odml.cache_update. Its BATCH_MATMUL(adjY) attention products are now odml.runtime_bmm. A signature input param_tensor INT32[1,1,1,7] was added; the runtime fills start/end. A graph exported with litert-torch's --apply_gpu_composites has this same shape, and the previous file was exported without that flag. The attention mask is applied as an ADD on a [bk, g, T, C] view of the logits, with the FLOAT32 mask broadcast; the mask input stays FLOAT32. A 16-token prefill signature, prefill_16, was added as a copy of prefill_128 that shares every weight buffer. Every other bundle section and the graph's weight region are byte-identical to the previous file (checked per section and per buffer). The new file is 2,059,236,272 bytes, sha256 7885496ccd8f70680474e66e88eccfdc966c1870a39b5fdbea837f897e4742a0. The previous file was 2,057,106,352 bytes, sha256 ea1838a7f9307815be7a28f86ef182d007d2d47423787e75584a4d48d541aa98. The checks compare three files: the previous file, an intermediate file with the rewrite but without prefill_16, and the new file. With fp32 activations on the Mac GPU, answers are byte-identical from the previous file to the intermediate (8/8 questions) and from the intermediate to the new file (9/9 prompts). On the CPU, the new file's answers match the intermediate's byte for byte (9/9 prompts), and prefill_16 gives bit-identical output to prefill_128 on the same tokens (2/2 cases). With the default fp16 activations, one long prompt got a wrong final answer from the new file and a right one from the intermediate. The speed rows below ran both files in the same window, with the Mac protocol under Performance. The Raspberry Pi 5, Pixel 8a and iPhone 17 Pro rows further down were measured on the previous file.
| Mac Studio M4 Max, GPU, fp16 activations | previous file | new file |
|---|---|---|
| 8 questions, correct final answers | 8/8 | 8/8 |
| GSM8K, 30 questions (max tokens 2048, greedy) | 29/30 | 30/30 |
| Prefill, 256-token prompt | 1376 tok/s | 1459 tok/s |
| Decode, 256 tokens | 92.0 tok/s | 125.6 tok/s |
| TTFT, 16-token prompt | 0.107 s | 0.035 s |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
β οΈ It's a reasoning model β give it room to think
VibeThinker solves with a step-by-step chain-of-thought, then a \boxed{} answer. Run it with
max_tokens β₯ 2048 β at a short limit it gets cut off before the answer. (All quality numbers
below were measured at 2048.)
Performance
Measured on the new file (see the 2026-10-06 note).
Apple M4 Max (macOS)
Mac Studio M4 Max, litert-lm 0.17.1 CLI, benchmark --cache no, 256 prompt and 256 decode tokens. GPU = WebGPU on Metal with fp16 activations, median of 3 processes. CPU = XNNPACK, median of 2 processes. The TTFT (16-token prompt) column comes from separate runs with a 16-token prompt and 32 decode tokens.
| Backend | Prefill (256) | Decode | TTFT | TTFT (16-token prompt) | Init |
|---|---|---|---|---|---|
| GPU (WebGPU on Metal) | 1459 tok/s | 125.6 tok/s | 0.19 s | 0.035 s | 2.8 s |
| CPU (XNNPACK) | 139 tok/s | 29.8 tok/s | 2.22 s | 0.526 s | 2.7 s |
Galaxy S26 (Android)
Samsung Galaxy S26 (SM-S942Q, Android 16), litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18. GPU = OpenCL delegate with the bundle's default fp16 activations, cold init (--disable_cache=true), 3 iterations in one process. CPU = XNNPACK with 4 threads and the XNNPACK weight cache, 2 iterations in one process. Each range runs from the lowest to the highest iteration in one process; the first GPU iteration follows the cold init. Peak is the process high-water mark (VmHWM).
| Backend | Prompt / decode tokens | Prefill | Decode | TTFT | Init | Peak (VmHWM) |
|---|---|---|---|---|---|---|
| GPU (OpenCL) | 1024 / 256 | 387.0β389.1 tok/s | 24.24β25.29 tok/s | 2.67β2.69 s | 12.75 s | 1.14 GB |
| GPU (OpenCL) | 16 / 32 | 232.3β252.1 tok/s | 19.57β23.88 tok/s | 0.11β0.12 s | 12.89 s | 1.14 GB |
| CPU (XNNPACK, 4 threads) | 1024 / 256 | 64.1β84.9 tok/s | 15.68β16.51 tok/s | 12.13β16.03 s | 4.11 s | 2.79 GB |
Accuracy note
Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought, max_tokens 2048, identical prompt and answer-extraction for every row).
| Configuration | GSM8K |
|---|---|
| bf16 (reference) | 97.0% |
| LiteRT int4 β block 32 | 90.0% (β7 pt) |
int4 (block 32) is at parity (β7 pt) and still 90% β strong for an on-device math model. bf16's 97% reflects this model's math specialization.
Why block 32 (not block 128)? This is a precision-sensitive math model: the coarser block-128 int4 (ΒΌ the dequant scales) collapsed to 64% (β33 pt) on GSM8K, while block 32 holds at 90%. So only the block-32 build is published. (Note: for general-purpose 4B reasoning models the opposite holds β block 128 is fine and faster β but exact arithmetic needs the finer block-32 grid.)
Galaxy S26 β GPU backend
The new file runs on the Galaxy S26 GPU backend, and 5 of 5 gate prompts were answered correctly. The OpenCL GPU delegate (LITERT_CL) takes every node of each transformer signature, in one partition each:
| signature | nodes on LITERT_CL |
|---|---|
prefill_128 |
1636 / 1636 |
prefill_16 |
1636 / 1636 |
decode |
1487 / 1487 |
XNNPACK takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128. The Mac GPU (WebGPU) delegates the same node counts. Measured on a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18; speed is in the Performance tables above.
GPU wiring, including the Gallery import toggle: GPU guide.
Usage
# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
--model_path model.litertlm \
--backend gpu \
--input_prompt "A bat and a ball cost \$1.10. The bat costs \$1.00 more than the ball. How much is the ball?"
The .litertlm bundle carries the tokenizer and prompt template (Qwen2 ChatML β
<|im_start|>role\nβ¦<|im_end|>), so no separate tokenizer files are needed.
Run on Android
Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β no computer or
adbneeded. The manual steps below are only required on older builds or for sideloading a local file.
The official Google AI Edge Gallery app runs
.litertlm models on-device:
- Install a recent Gallery (package
com.google.ai.edge.gallery, 1.0.15+ supports.litertlm). - Download
model.litertlmand push it:adb push model.litertlm /sdcard/Download/ - In the app tap +, pick the file, choose the GPU backend, and raise the max-tokens setting (β₯2048).
- Chat β the bundle already carries the tokenizer and Qwen2 chat template.
Measured on an 8 GB phone (added 2026-08-17, previous file): driving litert_lm_main directly on a Pixel 8a (Tensor G3, Mali-G715, 8 GB RAM), the previous graph ran entirely on the OpenCL delegate β 1603/1603 nodes in the 128-token prefill graph and 1452/1452 in decode, zero rejected ops. The new file was not re-measured on this phone. Worth knowing when reading its answers: this is a math-specialised reasoning model, so on general-knowledge prompts it reasons at length and can settle on a wrong answer β identically on CPU and GPU, which is how you can tell it is the model's domain rather than the backend.
Run on desktop (LiteRT-LM CLI)
The same .litertlm bundle runs on macOS / Linux / Windows with the official
LiteRT-LM CLI β including as a
local OpenAI-compatible API server:
pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/VibeThinker-3B model.litertlm vibethinker-3b
litert-lm run vibethinker-3b # interactive chat in the terminal
litert-lm serve
Run on iPhone
Verified on iPhone 17 Pro (LiteRT-LM Swift runtime): the block-32 build (1.62 GiB section, under the iOS limit) loads and generates correct answers (previous file).
Conversion
Converted with the official litert-torch
converter β a standard Qwen2ForCausalLM, no custom graph code. Recipe: blockwise-32 int4 + OCTAV
(INT4 weights, block 32, symmetric, OCTAV optimal-clipping), embeddings INT8, KV cache 4096.
from litert_torch.generative.export_hf.export import export
export(
model="WeiboAI/VibeThinker-3B",
output_dir="out",
quantization_recipe="qwen3_int4_block32_octav.json", # blockwise-32 int4 + OCTAV, int8 embeddings
cache_length=4096,
externalize_embedder=True,
)
Graph rewrite (2026-10-06). The current file was made from the previous one (Hub revision 087512d2eba6) with two tools in tools/gpu_graph/ of hf-to-litertlm. gpu_graph_retrofit.py makes the graph changes listed in the 2026-10-06 note; every original weight buffer keeps its bytes. prefill_bucket_clone.py copies an existing prefill signature (prefill_128) at a new length (16); the copy shares every weight buffer, and every existing subgraph and SignatureDef keeps its index and bytes.
python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
python tools/gpu_graph/prefill_bucket_clone.py build retrofit.litertlm new.litertlm --lengths 16 --source prefill_128 --report bucket.json
previous.litertlm is the previous file, new.litertlm the result published here as model.litertlm; LITERT_LM_CLI names the litert-lm CLI the scripts use to unpack and pack the bundle.
A rebuild with these commands gives every bundle section byte-identical to this file (sha256 per section). Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.
2026-08-29 β default system prompt restored (weights unchanged)
The upstream chat template emits a default system turn whenever the caller sends no system message β for this model: You are a helpful assistant.. The converter's template probe renders the template with a system message already present, so that block was never seen and never reached the bundle: with no system message the model was running without the default system turn it was tuned with. The chat template in model.litertlm now emits the block exactly once when no system message is given. In model.litertlm, the block is not emitted when you pass a system message. The restored block adds 11 prefill tokens to a conversation that sends no system message, so time-to-first-token grows by that much; per-token speed is unchanged.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file per token β only the file's own sha256 differs. What changed is the input: with no system message, the prompt now renders byte-identical to the upstream chat template's output, verified on the LiteRT-LM runtime. A system message you pass yourself renders as before. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-29, re-download.
Raspberry Pi 5 (CPU) (previous file)
Measured on the previous file (sha256 ea1838a7f9307815be7a28f86ef182d007d2d47423787e75584a4d48d541aa98) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical and the new file's Mac CPU decode and prefill are Γ1.037 and Γ0.985 of the previous file's, so the throughput columns are expected to hold.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (minβmax in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
model.litertlm |
22.5 (22.4β23.1) | 3.4 (3.4β3.4) | 13.6 s | 2.7 GB |
License
MIT, inherited from the base model WeiboAI/VibeThinker-3B.
- Downloads last month
- 1,461