Instructions to use litert-community/SmolLM3-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/SmolLM3-3B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/SmolLM3-3B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/SmolLM3-3B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).
Measured on device (edge-compat, smollm3-3b): Galaxy S26 Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 14.8 tok/s Β· prefill 268 tok/s Β· TTFT 820 ms Β· all 1476 ops delegated (2026-08-24); Pixel 8a Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 7.7 tok/s Β· prefill 12 tok/s Β· TTFT 1.56 s Β· all 1476 ops delegated (2026-08-17); Galaxy S26 Β· LiteRT-LM 0.16.0 Β· CPU Β· decode 9.8 tok/s Β· prefill 148 tok/s Β· TTFT 3.09 s (2026-09-05). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/smollm3-3b/CARD.md
Measured on device (edge-compat, smollm3-3b-int8): Galaxy S26 Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 14.4 tok/s Β· prefill 312 tok/s Β· TTFT 750 ms Β· all 1548 ops delegated (2026-08-24); Raspberry Pi 5 Β· LiteRT-LM 0.16.1 Β· CPU, 4 threads Β· decode 2.0 tok/s Β· prefill 19 tok/s Β· TTFT 14.04 s (2026-09-01). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/smollm3-3b-int8/CARD.md
SmolLM3-3B β LiteRT-LM (blockwise int4)
HuggingFaceTB/SmolLM3-3B
converted to the LiteRT-LM (.litertlm) format for on-device inference with
Google's LiteRT-LM runtime (the
engine behind the official litert-community/* models).
SmolLM3 is a fully-open 3B decoder (Apache-2.0) with GQA, a NoPE attention schedule, multilingual support, and long-context training β a strong small reasoner.
| File | SmolLM3-3B_q4_block32_ekv4096.litertlm (~2.0 GB) |
| Quantization | int4 weights β blockwise (block 32) + OCTAV optimal-clipping, symmetric; embedding INT8 |
| Compute | integer |
| Context (KV cache) | 4096 |
| Second file | SmolLM3-3B.litertlm (~3.1 GB), int8 |
| Base model | HuggingFaceTB/SmolLM3-3B |
2026-10-06: GPU graph rewrite of both files; the weights are unchanged. The previous files were exported without litert-torch's --apply_gpu_composites. Their graph wrote the KV cache with two DYNAMIC_UPDATE_SLICE ops per layer and ran attention with BATCH_MATMUL. The new graph has the shape that flag exports: one odml.cache_update composite per layer writes the KV cache, odml.runtime_bmm composites compute the attention products, and a signature input param_tensor (INT32[1,1,1,7]) is added for the runtime to fill. The attention mask stays FLOAT32 and is added to the logits by broadcast. Each file also gains a 16-token prefill signature, copied from its existing prefill signature. The copy shares every weight buffer, so the file grows by graph structure only. In each file, every other bundle section and every weight byte is identical to the previous file (checked per section and per buffer). SmolLM3-3B_q4_block32_ekv4096.litertlm (int4) is now 2,004,633,520 bytes, sha256 bdd7eb05b6764fb5a5173bdbdd59354a26f66e4c20b040f0d3da3532071d18f6 (previous: 2,002,257,840 bytes, sha256 119fc1f2β¦). SmolLM3-3B.litertlm (int8) is now 3,125,204,208 bytes, sha256 c1c6c635194745456d03ba3a851b9ebf3a2714218ce35d930d9eb93f7cf97e20 (previous: 3,109,328,112 bytes, sha256 f9aa844cβ¦). On the CPU, the 16-token signature's output is bit-identical to the existing prefill signature's on the same tokens (2 of 2 cases). The answers to 8 questions and 1 long prompt are byte-identical with and without the new signature, both on the CPU and on the Mac GPU with fp32 activations. With the default fp16 activations on the Mac GPU, the int4 file gets 7 of the 8 final answers right (6 without the new signature) and fails the long prompt in both forms. On 30 GSM8K questions (greedy, up to 2048 tokens) it gets 26 right, the same 26 as the previous file. The int8 file gets 7 of the 8 right and passes the long prompt, in both forms. The Raspberry Pi 5 and Pixel 8a figures further down were measured on the previous files and not re-measured.
Apple M4 Max, litert-lm 0.17.1, benchmark --cache no |
int4 file: previous β new | int8 file: previous β new |
|---|---|---|
| GPU decode, 256 tokens | 90.5 β 131.8 tok/s (Γ1.457) | 77.0 β 106.0 tok/s (Γ1.376) |
| GPU prefill, 256 tokens | 1344 β 1481 tok/s (Γ1.103) | 1414 β 1641 tok/s (Γ1.160) |
| GPU TTFT, 16-token prompt | 0.111 β 0.035 s (Γ0.312) | 0.205 β 0.038 s (Γ0.184) |
| CPU decode, 256 tokens | 24.8 β 24.5 tok/s (Γ0.990) | 24.4 β 24.2 tok/s (Γ0.990) |
| CPU TTFT, 16-token prompt | 1.325 β 0.542 s (Γ0.409) | 2.033 β 1.546 s (Γ0.761) |
| CPU init | 2.0 β 2.8 s (Γ1.451) | 6.7 β 9.4 s (Γ1.410) |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
Usage
Run with the LiteRT-LM runtime:
# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
--model_path SmolLM3-3B_q4_block32_ekv4096.litertlm \
--backend gpu \
--input_prompt "Explain on-device AI in one sentence."
The .litertlm bundle carries the tokenizer and the prompt template (ChatML β
<|im_start|>role / <|im_end|>, stop token <|im_end|>), so no separate
tokenizer files are needed.
Run on Android
Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β no computer or
adbneeded. The manual steps below are only required on older builds or for sideloading a local file.
The easiest way to try this model on a phone is the official
Google AI Edge Gallery app β it
runs .litertlm models fully on-device and can import your own:
- Install a recent Gallery (package
com.google.ai.edge.gallery, APK from the repo's releases β 1.0.15+ supports.litertlm). Older 1.0.x builds (packagecom.google.aiedge.gallery) only accept the legacy MediaPipe.taskformat and reject.litertlm. - Download
SmolLM3-3B_q4_block32_ekv4096.litertlmfrom this repo and push it to the device:adb push SmolLM3-3B_q4_block32_ekv4096.litertlm /sdcard/Download/ - In the app, tap the + button (bottom-right), pick the file, and choose the GPU backend (CPU also works).
- Chat. Nothing else to configure β the
.litertlmbundle already carries the tokenizer and ChatML prompt template.
See the Gallery
Importing Local Models
guide for details. To embed the model in your own Android app instead, use the
LiteRT-LM Kotlin API (Gradle artifact com.google.ai.edge.litertlm:litertlm-android,
getting started).
Measured on an 8 GB phone (added 2026-08-17, previous file): driving litert_lm_main directly on a Pixel 8a (Tensor G3, Mali-G715, 8 GB RAM), the previous graph ran entirely on the OpenCL delegate β 1476/1476 nodes in the 128-token prefill graph and 1308/1308 in decode, zero rejected ops β and answered correctly. The new file was not re-measured on this phone.
Run on desktop (LiteRT-LM CLI)
The same .litertlm bundle runs on macOS / Linux / Windows with the official
LiteRT-LM CLI β including as a
local OpenAI-compatible API server:
pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/SmolLM3-3B SmolLM3-3B_q4_block32_ekv4096.litertlm smollm3-3b
litert-lm run smollm3-3b # interactive chat in the terminal
litert-lm serve
Performance
The three speed tables below were measured on the new files (the 2026-10-06 upload).
Apple M4 Max
Mac Studio (M4 Max), litert-lm 0.17.1 (pip CLI), litert-lm benchmark --cache no: 256 prompt and 256 decode tokens, median of 3 processes on the GPU (WebGPU on Metal, fp16 activations) and of 2 on the CPU (XNNPACK). TTFT (16) comes from a separate run with a 16-token prompt and 32 decode tokens, same flags.
| File | Backend | Prefill (256) | Decode | TTFT (256) | TTFT (16) | Init |
|---|---|---|---|---|---|---|
SmolLM3-3B_q4_block32_ekv4096.litertlm |
GPU | 1481 tok/s | 131.8 tok/s | 0.18 s | 0.035 s | 2.7 s |
SmolLM3-3B_q4_block32_ekv4096.litertlm |
CPU | 140 tok/s | 24.5 tok/s | 2.23 s | 0.542 s | 2.8 s |
SmolLM3-3B.litertlm |
GPU | 1641 tok/s | 106.0 tok/s | 0.17 s | 0.038 s | 2.1 s |
SmolLM3-3B.litertlm |
CPU | 460 tok/s | 24.2 tok/s | 2.00 s | 1.546 s | 9.4 s |
--cache no runs the CPU without the XNNPACK weight cache. Without it, SmolLM3-3B.litertlm repacks its weights for the extra 16-token signature: init 6.62 β 9.62 s and peak footprint 6.86 β 9.48 GB on a 16-token prompt (without β with the signature). With the disk weight cache (the app default), init goes 0.41 β 0.56 s, footprint 1.01 β 0.87 GB, and TTFT at 16 tokens 0.7 β 0.17 s.
Galaxy S26
Samsung Galaxy S26 (SM-S942Q, Android 16), litert_lm_advanced_main built from LiteRT-LM main (2026-09-18 build) with libLiteRtOpenClAccelerator.so. GPU: OpenCL, bundle default fp16 activations, --disable_cache=true (cold init). CPU: XNNPACK, 4 threads, with the weight cache (the init writes it). Prefill, Decode and TTFT (1024): a 1024-token prompt and 256 decode tokens, 3 iterations in one process on the GPU and 2 on the CPU. TTFT (16): a 16-token prompt and 32 decode tokens, 3 iterations in one process. A range is the lowest and highest iteration. In the 1024-token GPU runs, decode was lowest on the third iteration; TTFT (16) was highest on the first iteration. Peak is the process high-water mark (VmHWM) of the 1024-token run.
| File | Backend | Prefill (1024) | Decode | TTFT (1024) | TTFT (16) | Init | Peak |
|---|---|---|---|---|---|---|---|
SmolLM3-3B_q4_block32_ekv4096.litertlm |
GPU | 372.5β383.2 tok/s | 19.45β25.09 tok/s | 2.71β2.80 s | 0.10β0.12 s | 13.37 s | 1.22 GB |
SmolLM3-3B_q4_block32_ekv4096.litertlm |
CPU | 55.6β76.2 tok/s | 11.49 tok/s | 13.53β18.49 s | 0.22β1.23 s | 2.59 s | 3.12 GB |
SmolLM3-3B.litertlm |
GPU | 221.3β458.5 tok/s | 12.48β17.22 tok/s | 2.29β4.71 s | 0.13β0.15 s | 6.93 s | 1.00 GB |
SmolLM3-3B.litertlm |
CPU | 157.6β174.3 tok/s | 9.56β9.83 tok/s | 5.98β6.60 s | not measured | 7.04 s | 4.36 GB |
iPhone 18 Pro β SmolLM3-3B_q4_block32_ekv4096.litertlm
iPhone 18 Pro (iPhone19,2, iOS 27.0), LiteRT-LM v0.17.1 xcframework (same bytes as v0.17.0) through a self-built harness app (C API / Swift wrapper), default memory entitlement (~3,369 MB). GPU: Metal. CPU: XNNPACK with the disk weight cache. Each row is 2 timed runs, each after a 6-minute rest; the thermal state was nominal at the start and end of every run.
| Backend | Prompt | Prefill | Decode | TTFT | Init |
|---|---|---|---|---|---|
| GPU (Metal) | 128 tokens | 303.7 tok/s | 43.9 tok/s | 0.44 s | 6.24 s |
| GPU (Metal) | 16 tokens | 359.1 tok/s | 43.3 tok/s | 0.071 s | 6.26 s |
| CPU (XNNPACK) | 128 tokens | 114.5 tok/s | 22.5 tok/s | 2.00 s | 2.01 s |
| CPU (XNNPACK) | 16 tokens | 94.4 tok/s | 19.1 tok/s | 0.530 s | 2.07 s |
Accuracy note
Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought asking for #### <n>,
identical prompt and answer-extraction for both rows β only the quantization differs).
| Configuration | GSM8K |
|---|---|
| bf16 (reference) | 81.0% |
| This model β LiteRT int4 (BOCTAV4) | 81.0% |
LiteRT int4 is fully at parity β 0.0 pt vs the bf16 reference. The blockwise-32 +
OCTAV recipe with a 4096 KV cache preserves reasoning accuracy exactly at n=100. The
model produces visible step-by-step chain-of-thought in the answer body and
terminates cleanly at <|im_end|> (no rambling).
Galaxy S26 β GPU backend
On the Galaxy S26 GPU, every transformer signature of both files is fully delegated: one partition each on the OpenCL delegate (LITERT_CL), with the same node counts as on the Mac.
| file | existing prefill signature | prefill_16 |
decode |
|---|---|---|---|
SmolLM3-3B_q4_block32_ekv4096.litertlm |
prefill_128: 1509 / 1509 nodes |
1509 / 1509 | 1343 / 1343 |
SmolLM3-3B.litertlm |
prefill_256: 1512 / 1512 nodes |
1512 / 1512 | 1346 / 1346 |
SmolLM3-3B.litertlm answered 5 of 5 gate prompts correctly on this GPU. Measured on a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main built from LiteRT-LM main (2026-09-18 build) and libLiteRtOpenClAccelerator.so; GPU activations are the bundle default fp16. Speed and peak memory are in the Performance tables above.
GPU wiring, including the Gallery import toggle: GPU guide.
Conversion
Converted with litert-torch via its
generic export_hf path. SmolLM3ForCausalLM rides the existing converter with no
custom code: the NoPE attention schedule (rotary disabled on every 4th layer,
no_rope_layer_interval=4) lowers to generic ops with no custom kernel. The int4
recipe is blockwise (block 32) + OCTAV optimal-clipping with the embedding kept
at INT8; the embedding is externalized into its own bundle section so the main
weights section stays under the iOS ~2 GiB single-mmap limit. Blockwise (not
channelwise) int4 plus OCTAV is what holds reasoning accuracy at parity.
Graph rewrite (2026-10-06). Both files were rebuilt from the previous files with two tools from tools/gpu_graph/ in https://github.com/john-rocky/hf-to-litertlm. gpu_graph_retrofit.py turns the two DYNAMIC_UPDATE_SLICE KV-cache writes per layer into one odml.cache_update composite, turns the BATCH_MATMUL attention products into odml.runtime_bmm, adds the param_tensor input, adds the FLOAT32 attention mask to the logits by broadcast (--mask add_bcast), and keeps every original weight buffer's bytes. prefill_bucket_clone.py copies an existing prefill signature at a new length; the copy shares every weight buffer, and every existing subgraph and SignatureDef keeps its index and bytes.
python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
python tools/gpu_graph/prefill_bucket_clone.py build retrofit.litertlm new.litertlm --lengths 16 --source prefill_128 --report bucket.json
For SmolLM3-3B.litertlm, whose only prefill signature is prefill_256, the second command takes --source prefill_256. A rebuild from the previous files gives every bundle section byte-identical to the shipped files. Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.
Training data & PII
This is a weights-exact format conversion of HuggingFaceTB/SmolLM3-3B; no new training was performed. SmolLM3 was trained by Hugging Face on ~11T tokens of publicly documented data β web (FineWeb-Edu, DCLM), code (StarCoder-family), math, and multilingual sources β plus public SFT/preference sets. Being web-derived it may incidentally contain PII; none was deliberately collected and this format conversion adds none. Apply your own content/PII filtering before deployment. See the base model card for the full data mixture.
2026-08-28 β start_token fix (weights unchanged)
The bundle's LlmMetadata start_token held the literal string "None". This tokenizer has no BOS, and the LiteRT-LM engine resolved that string to a real vocabulary token β so every prompt began with the word None, which the model was never trained on. The start token has been removed.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file β only the file's own sha256 differs. What changed is the input: the token stream the model reads for a given conversation can differ from the previous file's, and it now matches this model's own reference chat stream. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file in either direction. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-28, re-download.
2026-08-29 β default system prompt restored (weights unchanged)
The upstream chat template emits a default system turn whenever the caller sends no system message β for this model: the ## Metadata header (knowledge cutoff, today's date, Reasoning Mode: /think) followed by the instructions You are a helpful AI assistant named SmolLM, trained by Hugging Faceβ¦. The converter's template probe renders the template with a system message already present, so that block was never seen and never reached the bundle: with no system message the model was running without the default system turn it was tuned with. The chat template in SmolLM3-3B.litertlm, SmolLM3-3B_q4_block32_ekv4096.litertlm now emits the block exactly once when no system message is given. In SmolLM3-3B.litertlm, a system message you pass is wrapped in the same metadata header as upstream, and /no_think in it switches reasoning off. In SmolLM3-3B_q4_block32_ekv4096.litertlm, the block is not emitted when you pass a system message; the upstream template also wraps a caller's system message in its own preamble, and this file passes it through unchanged, exactly as it did before. The previous SmolLM3-3B.litertlm also forced an empty <think></think> before every answer (/no_think); upstream's default is /think, and the template now follows it. The restored block adds 240 prefill tokens to a conversation that sends no system message, so time-to-first-token grows by that much; per-token speed is unchanged.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file per token β only the file's own sha256 differs. What changed is the input: with no system message, the prompt now renders byte-identical to the upstream chat template's output, verified on the LiteRT-LM runtime. A system message you pass yourself now renders through the upstream metadata header in SmolLM3-3B.litertlm; in SmolLM3-3B_q4_block32_ekv4096.litertlm it renders as before. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-29, re-download.
2026-08-30 β tokenizer section replaced (weights unchanged)
The tokenizer in SmolLM3-3B.litertlm was a SentencePiece conversion of the model's BPE tokenizer, and the conversion lost the byte-level semantics: a standalone accented letter or symbol (Γ©, Γ±, ΓΌ, Β°, Β·, β¦) was encoded to the id of a single-byte token instead of the token the upstream tokenizer uses, and any character without a whole-character vocabulary entry (emoji, most of Latin Extended-A) became the token the conversion had reused as UNK β the end-of-text token for this vocabulary. For this file that token was <|im_end|>, so every turn's end marker arrived at the model as six spelled-out tokens instead of the single <|im_end|> id, and an emoji in a message arrived as an end-of-turn. SmolLM3-3B.litertlm now embeds the upstream tokenizer.json (the same HF tokenizer path most bundles in this collection use).
Tokenizer-only change: every section of the bundle except the tokenizer is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the chat template are unchanged and the speed and memory numbers on this card still describe this file per token β only the file's own sha256 differs. Verified on the LiteRT-LM runtime: the default turn, 7 probe strings, the 223 standalone characters U+00A1βU+017F and every special token now tokenize identically to the upstream tokenizer, and the four ASCII-only test questions answer byte-identically to the previous file (same ids in, same tokens out). Prompts containing accented letters, symbols or emoji reach the model differently from before, so individual answers to such prompts can change. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file, and any on-device rows were measured on the previous file too β the on-device gate has not been re-run on this one (the runtime's tokenizer code is the same on macOS and on device; the weights and graph are byte-identical). If you downloaded before 2026-08-30, re-download.
2026-08-31 β thought channel declared (metadata only, weights unchanged)
SmolLM3-3B.litertlm, SmolLM3-3B_q4_block32_ekv4096.litertlm now declare the reasoning channel in their metadata (LlmMetadata.channels: channel name thought, markers <think>β¦</think> exactly as this model emits them). Without the declaration the runtime has no way to tell the reasoning apart from the answer: the raw thinking streamed inline into the visible text, and a thinking_token_budget was silently ignored (the API returns OK and only logs a warning). With the channel declared, LiteRT-LM returns the reasoning separated in channels["thought"] and the thinking budget takes effect.
Metadata-only change: every section of the bundle except LlmMetadata is byte-identical to the previous file (verified per section, tokenizer included), so the weights, the graph, the tokenizer and the chat template are unchanged and the speed and accuracy numbers on this card still describe this file β only the file's own sha256 differs. Verified on the LiteRT-LM runtime (litert-lm-api 0.16.1): the visible answer stays clean, the reasoning lands in channels["thought"], and on one bundle of this batch thinking_token_budget=16 was confirmed to truncate the reasoning at exactly 16 tokens where it was a no-op before. If you downloaded before 2026-08-31, re-download to get the channel-aware file.
Raspberry Pi 5 (CPU) (previous file)
These rows were measured before the 2026-10-06 graph rewrite, on the previous graph (the files it replaced: SmolLM3-3B.litertlm sha256 f9aa844cβ¦, SmolLM3-3B_q4_block32_ekv4096.litertlm sha256 119fc1f2β¦), and were not re-measured; the weights are byte-identical and the new files' Mac CPU prefill, decode and TTFT are Γ0.982βΓ1.004 of the previous files', so the speed columns are expected to hold.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (minβmax in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
SmolLM3-3B.litertlm |
19.1 (18.4β20.4) | 2.0 (2.0β2.0) | 14.0 s | 4.1 GB |
SmolLM3-3B_q4_block32_ekv4096.litertlm |
18.3 (18.3β18.4) | 2.5 (2.5β2.5) | 16.3 s | 3.1 GB |
License
Apache-2.0, inherited from the base model HuggingFaceTB/SmolLM3-3B.
- Downloads last month
- 2,440
Model tree for litert-community/SmolLM3-3B
Base model
HuggingFaceTB/SmolLM3-3B-Base