Instructions to use litert-community/MiniCPM5-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/MiniCPM5-2B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Update README.md
#2
by deleted - opened
README.md
CHANGED
|
@@ -5,6 +5,7 @@ language:
|
|
| 5 |
- zh
|
| 6 |
base_model:
|
| 7 |
- openbmb/MiniCPM5-2B
|
|
|
|
| 8 |
pipeline_tag: text-generation
|
| 9 |
library_name: litert
|
| 10 |
tags:
|
|
@@ -18,14 +19,25 @@ tags:
|
|
| 18 |
|
| 19 |
# MiniCPM5-2B (LiteRT-LM)
|
| 20 |
|
| 21 |
-
This repository hosts the [**LiteRT-LM**](https://ai.google.dev/edge/litert-lm) (LiteRT formerly known as TensorFlow Lite) version of **MiniCPM5-2B**, optimized for fully on-device inference on mobile and edge hardware.
|
| 22 |
|
| 23 |
---
|
| 24 |
|
| 25 |
## Available Models
|
| 26 |
|
| 27 |
-
|
| 28 |
-
*
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
## What is MiniCPM?
|
| 31 |
|
|
@@ -73,7 +85,7 @@ Install `uv` and run the model directly from the LiteRT-LM command line:
|
|
| 73 |
|
| 74 |
```bash
|
| 75 |
uv tool install litert-lm
|
| 76 |
-
uvx litert-lm run --from-huggingface-repo=litert-community/MiniCPM5-2B
|
| 77 |
```
|
| 78 |
|
| 79 |
## Links
|
|
|
|
| 5 |
- zh
|
| 6 |
base_model:
|
| 7 |
- openbmb/MiniCPM5-2B
|
| 8 |
+
base_model_relation: quantized
|
| 9 |
pipeline_tag: text-generation
|
| 10 |
library_name: litert
|
| 11 |
tags:
|
|
|
|
| 19 |
|
| 20 |
# MiniCPM5-2B (LiteRT-LM)
|
| 21 |
|
| 22 |
+
This repository hosts the [**LiteRT-LM**](https://ai.google.dev/edge/litert-lm) (LiteRT formerly known as TensorFlow Lite) version of **[openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)**, optimized for fully on-device inference on mobile and edge hardware.
|
| 23 |
|
| 24 |
---
|
| 25 |
|
| 26 |
## Available Models
|
| 27 |
|
| 28 |
+
Recommend to use below both **CPU and GPU compatable** models, originally from [mlboydaisuke/MiniCPM5-2B-LiteRT](https://huggingface.co/mlboydaisuke/MiniCPM5-2B-LiteRT)
|
| 29 |
+
**Requires litert-lm ≥ 0.16** (thought channel + `ThinkingConfig`); measured here on 0.17.0.
|
| 30 |
+
|
| 31 |
+
| File | Recipe | Size |
|
| 32 |
+
|---|---|---|
|
| 33 |
+
| **`MiniCPM5-2B_int4.litertlm`** | int4 blockwise-32 + OCTAV on linears, int8 embedding | 1.55 GB |
|
| 34 |
+
| **`MiniCPM5-2B_int8.litertlm`** | int8 dynamic on linears + embedding; **fp32 activations declared** (see notes) | 2.60 GB |
|
| 35 |
+
|
| 36 |
+
The **int4 file is the phone file** (smaller, fastest GPU decode on every platform measured) — best used for direct answers or short reasoning; see the thinking-mode note below. **int8 is the file for reasoning that has to complete**: its thinking chains are ~3–4× shorter than int4's on the same questions and terminate where int4 runs into the token budget. int8's main weight section is 2.33 GB, above the single-section mmap ceiling of default-entitlement iOS apps, so it is a desktop / Android build.
|
| 37 |
+
|
| 38 |
+
Further, below are some CPU-only models for exploration:
|
| 39 |
+
* `minicpm_wi4c_wi8_afp32.litertlm`: This model features mixed INT4/INT8 weight-only quantization with FP32 activations (afp32). MLP projections use channelwise INT4 with Hadamard rotation; all remaining weights (attention, embedding, and lmhead) use channelwise INT8.
|
| 40 |
+
* `minicpm_wi8_afp32.litertlm`: This model features weight-only INT8 quantization (wi8) with FP32 activations (afp32).
|
| 41 |
|
| 42 |
## What is MiniCPM?
|
| 43 |
|
|
|
|
| 85 |
|
| 86 |
```bash
|
| 87 |
uv tool install litert-lm
|
| 88 |
+
uvx litert-lm run --from-huggingface-repo=litert-community/MiniCPM5-2B MiniCPM5-2B_int4.litertlm --prompt="What is the capital of France?"
|
| 89 |
```
|
| 90 |
|
| 91 |
## Links
|