File size: 4,895 Bytes
ec7c1fd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
---
license: apache-2.0
base_model:
  - guoxuter/ov_intent_analysis_sft
tags:
  - rockchip
  - rk3588
  - rkllm
  - npu
  - openviking
  - retrieval
  - intent-analysis
  - query-planning
  - qwen3.5
pipeline_tag: text-generation
---

# ov_intent_analysis_sft β€” RKLLM (RK3588) conversion

Pre-converted **W8A8** RKLLM runtime file of [`guoxuter/ov_intent_analysis_sft`](https://huggingface.co/guoxuter/ov_intent_analysis_sft) (the OpenViking retrieval intent-analysis / query-planner model, fine-tuned from Qwen3.5-0.8B).

Run the **same tuned query planner** on your RK3588 NPU β€” no x86 conversion rig required.

## Why this exists

OpenViking's docs recommend this model for the `query_planner` slot: it decides whether a search needs context retrieval, skips chitchat (no queries β†’ no token spend), and emits structured `skill` / `resource` / `memory` queries. The stock model ships as HF safetensors; to run it on the RK3588 NPU you need a `.rkllm` file, which can only be produced by **rkllm-toolkit (x86_64-only)**. This repo is that conversion, done once, so RK3588 owners can skip the whole rig.

## File

| File | Size | Spec |
|---|---|---|
| `ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm` | 1.3 GB | W8A8, RK3588, 3 NPU cores, max_context 4096 |

Runtime requirements: `librkllmrt.so` 1.3.0 (rkllm-toolkit 1.3.0 generation). Verified on kernel 6.1 vendor with rknpu driver 0.9.8.

## How to serve it

Any RKLLM-capable server works. Two options:

### Option A β€” full rkllama server (Ollama API; recommended for OpenViking)

```bash
pip install rkllama   # python 3.9–3.12
rkllama_server --models /path/to/models
```

Place the `.rkllm` in your models dir. This gives an Ollama-compatible API, which is what OpenViking's `query_planner` speaks.

### Option B β€” minimal OpenAI+Ollama server (no transformers/torch)

The same ctypes wrapper, no heavy deps (see the conversion recipe below for the gist). Serves `/v1/*` and `/api/*`.

## OpenViking wiring

```json
{
  "query_planner": {
    "provider": "litellm",
    "model": "ollama/guoxuter/ov_intent_analysis_sft:v7_q8",
    "api_base": "http://127.0.0.1:8091",
    "temperature": 0.0,
    "timeout": 60,
    "extra_request_body": { "think": false }
  }
}
```

**Keep the model string exactly as-is** β€” OpenViking auto-matches the bundled v7 prompt by string (`retrieval.ov_intent_analysis_sft_v7` in `intent_analyzer.py`). Only `api_base` changes: point it at your RKLLM server instead of Ollama.

## Benchmark (RK3588, same prompt)

| Runtime | Wall time | Notes |
|---|---|---|
| Ollama / CPU (GGUF Q8) | 16.5 s | output lands in `thinking` unless `think:false` |
| rk-llama.cpp NPU (GGUF Q8) | 10.9 s | needed `--reasoning off` |
| **RKLLM NPU (this file, W8A8)** | **~9 s** (1.8 s warm) | prefill ~200 t/s, decode ~13 t/s |

Query planning is prefill-dominated, which is exactly where the NPU wins. Decode is memory-bandwidth-bound, so don't expect magic on long generations β€” this model's outputs are short JSON.

## Conversion recipe (for reproducing / other models)

The converter is **x86-only**, so run it on an x86 box (any Linux, or a serverless cloud like Modal):

```python
# rkllm-toolkit 1.3.0 (wheel from airockchip/rknn-llm release-v1.3.0,
# rkllm-toolkit/packages/rkllm_toolkit-1.3.0-cp311-cp311-linux_x86_64.whl)
# deps pinned from that release's requirements.txt (torch 2.6.0, transformers 5.8.0, ...)

from rkllm.api import RKLLM

llm = RKLLM()
llm.load_huggingface(model="guoxuter/ov_intent_analysis_sft", device="cpu")  # or cuda
llm.build(
    do_quantization=True,
    optimization_level=1,
    quantized_dtype="W8A8",
    quantized_algorithm="normal",
    target_platform="RK3588",
    num_npu_core=3,
    dataset="data_quant.json",   # calibration: input/target pairs
    hybrid_rate=0,               # REQUIRED arg in 1.3.0
    max_context=4096,            # NOTE: `max_context`, NOT `max_context_len`
)
llm.export_rkllm("ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm")
```

Gotchas hit along the way (so you don't):
- `build()` takes `max_context`, not `max_context_len` (raises `TypeError: unexpected keyword argument`).
- `hybrid_rate=0` is required in the 1.3.0 signature.
- The toolkit only ships **x86_64 wheels** β€” there is no aarch64 path; don't fight it on an ARM SBC.
- Keep the toolkit major version aligned with your runtime's `librkllmrt.so` (1.3.0 ↔ 1.3.0).

## Attribution & license

- Base model: [`guoxuter/ov_intent_analysis_sft`](https://huggingface.co/guoxuter/ov_intent_analysis_sft) β€” **Apache-2.0**, which its card states covers the fine-tuned checkpoint (same license as Qwen3.5-0.8B).
- Conversion performed with Rockchip's rkllm-toolkit 1.3.0 ([airockchip/rknn-llm](https://github.com/airockchip/rknn-llm)).
- This conversion is published under **Apache-2.0**.

Big thanks to guoxuter for the tuned model and the OpenViking team for the recommended workflow.