n8dgr8 commited on
Commit
ec7c1fd
Β·
verified Β·
1 Parent(s): a2da734

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,119 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - guoxuter/ov_intent_analysis_sft
5
+ tags:
6
+ - rockchip
7
+ - rk3588
8
+ - rkllm
9
+ - npu
10
+ - openviking
11
+ - retrieval
12
+ - intent-analysis
13
+ - query-planning
14
+ - qwen3.5
15
+ pipeline_tag: text-generation
16
+ ---
17
+
18
+ # ov_intent_analysis_sft β€” RKLLM (RK3588) conversion
19
+
20
+ Pre-converted **W8A8** RKLLM runtime file of [`guoxuter/ov_intent_analysis_sft`](https://huggingface.co/guoxuter/ov_intent_analysis_sft) (the OpenViking retrieval intent-analysis / query-planner model, fine-tuned from Qwen3.5-0.8B).
21
+
22
+ Run the **same tuned query planner** on your RK3588 NPU β€” no x86 conversion rig required.
23
+
24
+ ## Why this exists
25
+
26
+ OpenViking's docs recommend this model for the `query_planner` slot: it decides whether a search needs context retrieval, skips chitchat (no queries β†’ no token spend), and emits structured `skill` / `resource` / `memory` queries. The stock model ships as HF safetensors; to run it on the RK3588 NPU you need a `.rkllm` file, which can only be produced by **rkllm-toolkit (x86_64-only)**. This repo is that conversion, done once, so RK3588 owners can skip the whole rig.
27
+
28
+ ## File
29
+
30
+ | File | Size | Spec |
31
+ |---|---|---|
32
+ | `ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm` | 1.3 GB | W8A8, RK3588, 3 NPU cores, max_context 4096 |
33
+
34
+ Runtime requirements: `librkllmrt.so` 1.3.0 (rkllm-toolkit 1.3.0 generation). Verified on kernel 6.1 vendor with rknpu driver 0.9.8.
35
+
36
+ ## How to serve it
37
+
38
+ Any RKLLM-capable server works. Two options:
39
+
40
+ ### Option A β€” full rkllama server (Ollama API; recommended for OpenViking)
41
+
42
+ ```bash
43
+ pip install rkllama # python 3.9–3.12
44
+ rkllama_server --models /path/to/models
45
+ ```
46
+
47
+ Place the `.rkllm` in your models dir. This gives an Ollama-compatible API, which is what OpenViking's `query_planner` speaks.
48
+
49
+ ### Option B β€” minimal OpenAI+Ollama server (no transformers/torch)
50
+
51
+ The same ctypes wrapper, no heavy deps (see the conversion recipe below for the gist). Serves `/v1/*` and `/api/*`.
52
+
53
+ ## OpenViking wiring
54
+
55
+ ```json
56
+ {
57
+ "query_planner": {
58
+ "provider": "litellm",
59
+ "model": "ollama/guoxuter/ov_intent_analysis_sft:v7_q8",
60
+ "api_base": "http://127.0.0.1:8091",
61
+ "temperature": 0.0,
62
+ "timeout": 60,
63
+ "extra_request_body": { "think": false }
64
+ }
65
+ }
66
+ ```
67
+
68
+ **Keep the model string exactly as-is** β€” OpenViking auto-matches the bundled v7 prompt by string (`retrieval.ov_intent_analysis_sft_v7` in `intent_analyzer.py`). Only `api_base` changes: point it at your RKLLM server instead of Ollama.
69
+
70
+ ## Benchmark (RK3588, same prompt)
71
+
72
+ | Runtime | Wall time | Notes |
73
+ |---|---|---|
74
+ | Ollama / CPU (GGUF Q8) | 16.5 s | output lands in `thinking` unless `think:false` |
75
+ | rk-llama.cpp NPU (GGUF Q8) | 10.9 s | needed `--reasoning off` |
76
+ | **RKLLM NPU (this file, W8A8)** | **~9 s** (1.8 s warm) | prefill ~200 t/s, decode ~13 t/s |
77
+
78
+ Query planning is prefill-dominated, which is exactly where the NPU wins. Decode is memory-bandwidth-bound, so don't expect magic on long generations β€” this model's outputs are short JSON.
79
+
80
+ ## Conversion recipe (for reproducing / other models)
81
+
82
+ The converter is **x86-only**, so run it on an x86 box (any Linux, or a serverless cloud like Modal):
83
+
84
+ ```python
85
+ # rkllm-toolkit 1.3.0 (wheel from airockchip/rknn-llm release-v1.3.0,
86
+ # rkllm-toolkit/packages/rkllm_toolkit-1.3.0-cp311-cp311-linux_x86_64.whl)
87
+ # deps pinned from that release's requirements.txt (torch 2.6.0, transformers 5.8.0, ...)
88
+
89
+ from rkllm.api import RKLLM
90
+
91
+ llm = RKLLM()
92
+ llm.load_huggingface(model="guoxuter/ov_intent_analysis_sft", device="cpu") # or cuda
93
+ llm.build(
94
+ do_quantization=True,
95
+ optimization_level=1,
96
+ quantized_dtype="W8A8",
97
+ quantized_algorithm="normal",
98
+ target_platform="RK3588",
99
+ num_npu_core=3,
100
+ dataset="data_quant.json", # calibration: input/target pairs
101
+ hybrid_rate=0, # REQUIRED arg in 1.3.0
102
+ max_context=4096, # NOTE: `max_context`, NOT `max_context_len`
103
+ )
104
+ llm.export_rkllm("ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm")
105
+ ```
106
+
107
+ Gotchas hit along the way (so you don't):
108
+ - `build()` takes `max_context`, not `max_context_len` (raises `TypeError: unexpected keyword argument`).
109
+ - `hybrid_rate=0` is required in the 1.3.0 signature.
110
+ - The toolkit only ships **x86_64 wheels** β€” there is no aarch64 path; don't fight it on an ARM SBC.
111
+ - Keep the toolkit major version aligned with your runtime's `librkllmrt.so` (1.3.0 ↔ 1.3.0).
112
+
113
+ ## Attribution & license
114
+
115
+ - Base model: [`guoxuter/ov_intent_analysis_sft`](https://huggingface.co/guoxuter/ov_intent_analysis_sft) β€” **Apache-2.0**, which its card states covers the fine-tuned checkpoint (same license as Qwen3.5-0.8B).
116
+ - Conversion performed with Rockchip's rkllm-toolkit 1.3.0 ([airockchip/rknn-llm](https://github.com/airockchip/rknn-llm)).
117
+ - This conversion is published under **Apache-2.0**.
118
+
119
+ Big thanks to guoxuter for the tuned model and the OpenViking team for the recommended workflow.
ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c2d5477021b55ced9d5228f564663128a34281a2c43e159fd64323cdbd09a710
3
+ size 1299987468