LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

Measured on device (edge-compat): Galaxy S26 Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 13.9 tok/s Β· prefill 157 tok/s Β· TTFT 1.36 s Β· all 1603 ops delegated (2026-08-24); Pixel 8a Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 7.6 tok/s Β· prefill 13 tok/s Β· TTFT 1.48 s Β· all 1603 ops delegated (2026-08-17); Mac Studio M4 Max Β· LiteRT-LM 0.14.0 Β· GPU Β· decode 93.2 tok/s Β· prefill 1359 tok/s Β· TTFT 199 ms (2026-07-23). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/vibethinker-3b/CARD.md

VibeThinker-3B β€” LiteRT-LM (blockwise int4)

WeiboAI/VibeThinker-3B converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime (the engine behind the official litert-community/* models).

VibeThinker-3B is a dense 3B math/reasoning model (Qwen2ForCausalLM, 36 layers) β€” it solves problems with an inline chain-of-thought and is strong at arithmetic and math word problems. Standard Qwen2 architecture, so it rides the existing converter and runtime directly.

File model.litertlm β€” int4 block 32 (~1.9 GB)
Quantization int4 weights (symmetric) + OCTAV optimal-clipping; embeddings INT8 (externalized section)
Compute integer
Context (KV cache) 4096
Base model WeiboAI/VibeThinker-3B

2026-10-06: GPU graph rewrite. Only the prefill/decode graph of model.litertlm changed. Its two DYNAMIC_UPDATE_SLICE KV-cache writes per layer are now one STABLEHLO_COMPOSITE odml.cache_update. Its BATCH_MATMUL(adjY) attention products are now odml.runtime_bmm. A signature input param_tensor INT32[1,1,1,7] was added; the runtime fills start/end. A graph exported with litert-torch's --apply_gpu_composites has this same shape, and the previous file was exported without that flag. The attention mask is applied as an ADD on a [bk, g, T, C] view of the logits, with the FLOAT32 mask broadcast; the mask input stays FLOAT32. A 16-token prefill signature, prefill_16, was added as a copy of prefill_128 that shares every weight buffer. Every other bundle section and the graph's weight region are byte-identical to the previous file (checked per section and per buffer). The new file is 2,059,236,272 bytes, sha256 7885496ccd8f70680474e66e88eccfdc966c1870a39b5fdbea837f897e4742a0. The previous file was 2,057,106,352 bytes, sha256 ea1838a7f9307815be7a28f86ef182d007d2d47423787e75584a4d48d541aa98. The checks compare three files: the previous file, an intermediate file with the rewrite but without prefill_16, and the new file. With fp32 activations on the Mac GPU, answers are byte-identical from the previous file to the intermediate (8/8 questions) and from the intermediate to the new file (9/9 prompts). On the CPU, the new file's answers match the intermediate's byte for byte (9/9 prompts), and prefill_16 gives bit-identical output to prefill_128 on the same tokens (2/2 cases). With the default fp16 activations, one long prompt got a wrong final answer from the new file and a right one from the intermediate. The speed rows below ran both files in the same window, with the Mac protocol under Performance. The Raspberry Pi 5, Pixel 8a and iPhone 17 Pro rows further down were measured on the previous file.

Mac Studio M4 Max, GPU, fp16 activations previous file new file
8 questions, correct final answers 8/8 8/8
GSM8K, 30 questions (max tokens 2048, greedy) 29/30 30/30
Prefill, 256-token prompt 1376 tok/s 1459 tok/s
Decode, 256 tokens 92.0 tok/s 125.6 tok/s
TTFT, 16-token prompt 0.107 s 0.035 s

2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.

⚠️ It's a reasoning model β€” give it room to think

VibeThinker solves with a step-by-step chain-of-thought, then a \boxed{} answer. Run it with max_tokens β‰₯ 2048 β€” at a short limit it gets cut off before the answer. (All quality numbers below were measured at 2048.)

Performance

Measured on the new file (see the 2026-10-06 note).

Apple M4 Max (macOS)

Mac Studio M4 Max, litert-lm 0.17.1 CLI, benchmark --cache no, 256 prompt and 256 decode tokens. GPU = WebGPU on Metal with fp16 activations, median of 3 processes. CPU = XNNPACK, median of 2 processes. The TTFT (16-token prompt) column comes from separate runs with a 16-token prompt and 32 decode tokens.

Backend Prefill (256) Decode TTFT TTFT (16-token prompt) Init
GPU (WebGPU on Metal) 1459 tok/s 125.6 tok/s 0.19 s 0.035 s 2.8 s
CPU (XNNPACK) 139 tok/s 29.8 tok/s 2.22 s 0.526 s 2.7 s

Galaxy S26 (Android)

Samsung Galaxy S26 (SM-S942Q, Android 16), litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18. GPU = OpenCL delegate with the bundle's default fp16 activations, cold init (--disable_cache=true), 3 iterations in one process. CPU = XNNPACK with 4 threads and the XNNPACK weight cache, 2 iterations in one process. Each range runs from the lowest to the highest iteration in one process; the first GPU iteration follows the cold init. Peak is the process high-water mark (VmHWM).

Backend Prompt / decode tokens Prefill Decode TTFT Init Peak (VmHWM)
GPU (OpenCL) 1024 / 256 387.0–389.1 tok/s 24.24–25.29 tok/s 2.67–2.69 s 12.75 s 1.14 GB
GPU (OpenCL) 16 / 32 232.3–252.1 tok/s 19.57–23.88 tok/s 0.11–0.12 s 12.89 s 1.14 GB
CPU (XNNPACK, 4 threads) 1024 / 256 64.1–84.9 tok/s 15.68–16.51 tok/s 12.13–16.03 s 4.11 s 2.79 GB

Accuracy note

Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought, max_tokens 2048, identical prompt and answer-extraction for every row).

Configuration GSM8K
bf16 (reference) 97.0%
LiteRT int4 β€” block 32 90.0% (βˆ’7 pt)

int4 (block 32) is at parity (βˆ’7 pt) and still 90% β€” strong for an on-device math model. bf16's 97% reflects this model's math specialization.

Why block 32 (not block 128)? This is a precision-sensitive math model: the coarser block-128 int4 (ΒΌ the dequant scales) collapsed to 64% (βˆ’33 pt) on GSM8K, while block 32 holds at 90%. So only the block-32 build is published. (Note: for general-purpose 4B reasoning models the opposite holds β€” block 128 is fine and faster β€” but exact arithmetic needs the finer block-32 grid.)

Galaxy S26 β€” GPU backend

The new file runs on the Galaxy S26 GPU backend, and 5 of 5 gate prompts were answered correctly. The OpenCL GPU delegate (LITERT_CL) takes every node of each transformer signature, in one partition each:

signature nodes on LITERT_CL
prefill_128 1636 / 1636
prefill_16 1636 / 1636
decode 1487 / 1487

XNNPACK takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128. The Mac GPU (WebGPU) delegates the same node counts. Measured on a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18; speed is in the Performance tables above.

GPU wiring, including the Gallery import toggle: GPU guide.

Usage

# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
  --model_path model.litertlm \
  --backend gpu \
  --input_prompt "A bat and a ball cost \$1.10. The bat costs \$1.00 more than the ball. How much is the ball?"

The .litertlm bundle carries the tokenizer and prompt template (Qwen2 ChatML β€” <|im_start|>role\n…<|im_end|>), so no separate tokenizer files are needed.

Run on Android

Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β€” no computer or adb needed. The manual steps below are only required on older builds or for sideloading a local file.

The official Google AI Edge Gallery app runs .litertlm models on-device:

  1. Install a recent Gallery (package com.google.ai.edge.gallery, 1.0.15+ supports .litertlm).
  2. Download model.litertlm and push it: adb push model.litertlm /sdcard/Download/
  3. In the app tap +, pick the file, choose the GPU backend, and raise the max-tokens setting (β‰₯2048).
  4. Chat β€” the bundle already carries the tokenizer and Qwen2 chat template.

Measured on an 8 GB phone (added 2026-08-17, previous file): driving litert_lm_main directly on a Pixel 8a (Tensor G3, Mali-G715, 8 GB RAM), the previous graph ran entirely on the OpenCL delegate β€” 1603/1603 nodes in the 128-token prefill graph and 1452/1452 in decode, zero rejected ops. The new file was not re-measured on this phone. Worth knowing when reading its answers: this is a math-specialised reasoning model, so on general-knowledge prompts it reasons at length and can settle on a wrong answer β€” identically on CPU and GPU, which is how you can tell it is the model's domain rather than the backend.

Run on desktop (LiteRT-LM CLI)

The same .litertlm bundle runs on macOS / Linux / Windows with the official LiteRT-LM CLI β€” including as a local OpenAI-compatible API server:

pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/VibeThinker-3B model.litertlm vibethinker-3b
litert-lm run vibethinker-3b     # interactive chat in the terminal
litert-lm serve

Run on iPhone

Verified on iPhone 17 Pro (LiteRT-LM Swift runtime): the block-32 build (1.62 GiB section, under the iOS limit) loads and generates correct answers (previous file).

Conversion

Converted with the official litert-torch converter β€” a standard Qwen2ForCausalLM, no custom graph code. Recipe: blockwise-32 int4 + OCTAV (INT4 weights, block 32, symmetric, OCTAV optimal-clipping), embeddings INT8, KV cache 4096.

from litert_torch.generative.export_hf.export import export
export(
    model="WeiboAI/VibeThinker-3B",
    output_dir="out",
    quantization_recipe="qwen3_int4_block32_octav.json",  # blockwise-32 int4 + OCTAV, int8 embeddings
    cache_length=4096,
    externalize_embedder=True,
)

Graph rewrite (2026-10-06). The current file was made from the previous one (Hub revision 087512d2eba6) with two tools in tools/gpu_graph/ of hf-to-litertlm. gpu_graph_retrofit.py makes the graph changes listed in the 2026-10-06 note; every original weight buffer keeps its bytes. prefill_bucket_clone.py copies an existing prefill signature (prefill_128) at a new length (16); the copy shares every weight buffer, and every existing subgraph and SignatureDef keeps its index and bytes.

python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
python tools/gpu_graph/prefill_bucket_clone.py build retrofit.litertlm new.litertlm --lengths 16 --source prefill_128 --report bucket.json

previous.litertlm is the previous file, new.litertlm the result published here as model.litertlm; LITERT_LM_CLI names the litert-lm CLI the scripts use to unpack and pack the bundle.

A rebuild with these commands gives every bundle section byte-identical to this file (sha256 per section). Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.

2026-08-29 β€” default system prompt restored (weights unchanged)

The upstream chat template emits a default system turn whenever the caller sends no system message β€” for this model: You are a helpful assistant.. The converter's template probe renders the template with a system message already present, so that block was never seen and never reached the bundle: with no system message the model was running without the default system turn it was tuned with. The chat template in model.litertlm now emits the block exactly once when no system message is given. In model.litertlm, the block is not emitted when you pass a system message. The restored block adds 11 prefill tokens to a conversation that sends no system message, so time-to-first-token grows by that much; per-token speed is unchanged.

Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file per token β€” only the file's own sha256 differs. What changed is the input: with no system message, the prompt now renders byte-identical to the upstream chat template's output, verified on the LiteRT-LM runtime. A system message you pass yourself renders as before. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-29, re-download.

Raspberry Pi 5 (CPU) (previous file)

Measured on the previous file (sha256 ea1838a7f9307815be7a28f86ef182d007d2d47423787e75584a4d48d541aa98) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical and the new file's Mac CPU decode and prefill are Γ—1.037 and Γ—0.985 of the previous file's, so the throughput columns are expected to hold.

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (min–max in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.

File Prefill (tok/s) Decode (tok/s) TTFT Peak RSS
model.litertlm 22.5 (22.4–23.1) 3.4 (3.4–3.4) 13.6 s 2.7 GB

License

MIT, inherited from the base model WeiboAI/VibeThinker-3B.

Downloads last month
1,461
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for litert-community/VibeThinker-3B

Base model

Qwen/Qwen2.5-3B
Quantized
(55)
this model

Collection including litert-community/VibeThinker-3B