Update README.md

#2
by deleted - opened
Files changed (1) hide show
  1. README.md +16 -4
README.md CHANGED
@@ -5,6 +5,7 @@ language:
5
  - zh
6
  base_model:
7
  - openbmb/MiniCPM5-2B
 
8
  pipeline_tag: text-generation
9
  library_name: litert
10
  tags:
@@ -18,14 +19,25 @@ tags:
18
 
19
  # MiniCPM5-2B (LiteRT-LM)
20
 
21
- This repository hosts the [**LiteRT-LM**](https://ai.google.dev/edge/litert-lm) (LiteRT formerly known as TensorFlow Lite) version of **MiniCPM5-2B**, optimized for fully on-device inference on mobile and edge hardware.
22
 
23
  ---
24
 
25
  ## Available Models
26
 
27
- * **`minicpm_wi4c_wi8_afp32.litertlm`**: This model features mixed INT4/INT8 weight-only quantization with FP32 activations (afp32). MLP projections use channelwise INT4 with Hadamard rotation; all remaining weights (attention, embedding, and lmhead) use channelwise INT8.
28
- * **`minicpm_wi8_afp32.litertlm`**: This model features weight-only INT8 quantization (wi8) with FP32 activations (afp32).
 
 
 
 
 
 
 
 
 
 
 
29
 
30
  ## What is MiniCPM?
31
 
@@ -73,7 +85,7 @@ Install `uv` and run the model directly from the LiteRT-LM command line:
73
 
74
  ```bash
75
  uv tool install litert-lm
76
- uvx litert-lm run --from-huggingface-repo=litert-community/MiniCPM5-2B minicpm_wi4c_wi8_afp32.litertlm --prompt="What is the capital of France?"
77
  ```
78
 
79
  ## Links
 
5
  - zh
6
  base_model:
7
  - openbmb/MiniCPM5-2B
8
+ base_model_relation: quantized
9
  pipeline_tag: text-generation
10
  library_name: litert
11
  tags:
 
19
 
20
  # MiniCPM5-2B (LiteRT-LM)
21
 
22
+ This repository hosts the [**LiteRT-LM**](https://ai.google.dev/edge/litert-lm) (LiteRT formerly known as TensorFlow Lite) version of **[openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)**, optimized for fully on-device inference on mobile and edge hardware.
23
 
24
  ---
25
 
26
  ## Available Models
27
 
28
+ Recommend to use below both **CPU and GPU compatable** models, originally from [mlboydaisuke/MiniCPM5-2B-LiteRT](https://huggingface.co/mlboydaisuke/MiniCPM5-2B-LiteRT)
29
+ **Requires litert-lm ≥ 0.16** (thought channel + `ThinkingConfig`); measured here on 0.17.0.
30
+
31
+ | File | Recipe | Size |
32
+ |---|---|---|
33
+ | **`MiniCPM5-2B_int4.litertlm`** | int4 blockwise-32 + OCTAV on linears, int8 embedding | 1.55 GB |
34
+ | **`MiniCPM5-2B_int8.litertlm`** | int8 dynamic on linears + embedding; **fp32 activations declared** (see notes) | 2.60 GB |
35
+
36
+ The **int4 file is the phone file** (smaller, fastest GPU decode on every platform measured) — best used for direct answers or short reasoning; see the thinking-mode note below. **int8 is the file for reasoning that has to complete**: its thinking chains are ~3–4× shorter than int4's on the same questions and terminate where int4 runs into the token budget. int8's main weight section is 2.33 GB, above the single-section mmap ceiling of default-entitlement iOS apps, so it is a desktop / Android build.
37
+
38
+ Further, below are some CPU-only models for exploration:
39
+ * `minicpm_wi4c_wi8_afp32.litertlm`: This model features mixed INT4/INT8 weight-only quantization with FP32 activations (afp32). MLP projections use channelwise INT4 with Hadamard rotation; all remaining weights (attention, embedding, and lmhead) use channelwise INT8.
40
+ * `minicpm_wi8_afp32.litertlm`: This model features weight-only INT8 quantization (wi8) with FP32 activations (afp32).
41
 
42
  ## What is MiniCPM?
43
 
 
85
 
86
  ```bash
87
  uv tool install litert-lm
88
+ uvx litert-lm run --from-huggingface-repo=litert-community/MiniCPM5-2B MiniCPM5-2B_int4.litertlm --prompt="What is the capital of France?"
89
  ```
90
 
91
  ## Links