Text Classification
Transformers
Safetensors
English
Chinese
qwen3_5
image-text-to-text
decision-model
system-one
decision-index
lora-merged
Instructions to use PelaAI/KnowLine-4B-Gen3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen3")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen3") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download INFERENCE.md from PelaAI/KnowLine-4B-Gen3: direct link, hf CLI and curl.
- Browser
- Download file 7.25 kB
-
https://huggingface.co/PelaAI/KnowLine-4B-Gen3/resolve/main/INFERENCE.md
- Command line
-
hf download hf://PelaAI/KnowLine-4B-Gen3/INFERENCE.md
-
curl -L -o INFERENCE.md https://huggingface.co/PelaAI/KnowLine-4B-Gen3/resolve/main/INFERENCE.md
7.25 kB
| # KnowLine-4B-Gen3: serving and Decision Index reproduction | |
| These settings reproduce our self-run Decision Index 0.3 evaluation of these weights: the full 0.3 public suite sent | |
| fresh (no rows carried over from another run), scored 2026-10-08 10:42 CST. | |
| ## Weights | |
| - This repository holds bf16 weights made by averaging merged models tensor by tensor (each step computed in fp32 and | |
| stored in bf16): | |
| - [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2), 0.5; | |
| - a model from the next training round (Qwen3.5-4B plus a LoRA r32, merged), 0.5. | |
| - 248 tensors are averaged. The other 490 are identical in both models and copied unchanged, including the base | |
| model's vision tower and MTP head. | |
| - The architecture is `Qwen3_5ForConditionalGeneration`. `config.json`, the tokenizer files and the chat template are | |
| byte-identical to Gen1 and Gen2. | |
| - `SHA256SUMS` lists every file. | |
| ## Software | |
| | component | version | | |
| |---|---| | |
| | SGLang | 0.5.21 (torch 2.13.0, CUDA 13.0, flashinfer-python 0.6.18) | | |
| | transformers | 5.12.1 | | |
| | front end | `knowline_server.py` (this repository; needs only transformers and requests) | | |
| | decision-index kit | 0.3 at commit 62d2f51de34a2de64906345b6bc3e98e27ff55c7 | | |
| ### About the front end | |
| `knowline_server.py` is the `/v1/systemone` front end of our runs, packaged as one file. It is the same file as in | |
| Gen2; only the model name in its docstring changed. It contains: | |
| - rendering: the chat template with thinking off; the state is followed by one user turn with the instruction, the | |
| question and all options labelled A, B, ...; the assistant turn opens with `Answer:`; | |
| - label-token scoring: one prefill per question, then a softmax over the option labels only; | |
| - the HTTP server. | |
| The Gen3 run was served by the same in-house front end as the Gen1 and Gen2 runs. Before the Gen3 run it received the | |
| same 2026-10-08 fix as `knowline_server.py` (see [Changes](#changes)), which gives the same answers on every request | |
| that worked before. `knowline_server.py` was checked against the in-house front end on Gen1: | |
| - prompts, answer keys and label token ids were identical on 7,500 Decision Index rows and training requests; | |
| - through a live SGLang engine, choices agreed on all 953 questions of 200 Decision Index rows. | |
| ## Serve | |
| `serve_knowline.sh` starts both processes. | |
| 1. **SGLang engine.** | |
| - FP8 is on at serving: weights are stored in bf16 and SGLang quantizes them on load. | |
| - `--mem-fraction-static 0.72` only sizes SGLang's KV cache. Any value works and scores do not depend on it. | |
| - SGLang treats the model as multimodal by itself. The Decision Index run sent text only. | |
| 2. **`/v1/systemone` front end:** `knowline_server.py`, `chat` style, temperature 1, 16 scoring threads, no per-type | |
| calibration file. | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests # torch / CUDA per SGLang's install docs | |
| bash serve_knowline.sh PelaAI/KnowLine-4B-Gen3 0 8080 # GPU 0; SGLang on :9080, /v1/systemone on :8080 | |
| curl -s http://127.0.0.1:8080/health | |
| curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{ | |
| "model": "m", | |
| "state": "Customer: my order arrived broken, I want my money back.", | |
| "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"}, | |
| "tone": {"type": "choice", "instructions": "Customer tone?", | |
| "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}' | |
| ``` | |
| Equivalent manual launch, with the exact flags of our run: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server --model-path PelaAI/KnowLine-4B-Gen3 --served-model-name m --tp 1 \ | |
| --quantization fp8 --mem-fraction-static 0.72 --mamba-radix-cache-strategy extra_buffer --enable-fp32-lm-head --port 9080 & | |
| python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend sglang --url http://127.0.0.1:9080 \ | |
| --served-model-name m --temperature 1 --workers 16 --port 8080 | |
| ``` | |
| Without SGLang (CPU or a single GPU through transformers; slower, no prefix cache, bf16 rather than FP8): | |
| ```bash | |
| pip install torch "transformers==5.12.1" requests | |
| python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend hf --port 8080 | |
| ``` | |
| ## Decision Index run | |
| - Suite rows: the 140,620 rows that `decision-index run --edition 0.3` sends: `selected-rows.jsonl.gz`, | |
| `added-rows.jsonl.gz` and `gsm8k-rows.jsonl.gz` of `suite-0.3`, with ToolRet, BRIGHT and ACOS cut to their scoring | |
| subsets. | |
| - Four identical servers ran on four GPUs, one each, with the settings above. The rows are split round-robin into 64 | |
| shards, 16 per server. Each shard is a resumable `decision-index run` against its server: | |
| ```bash | |
| decision-index run --edition 0.3 --engine http --option base_url=http://127.0.0.1:<port> --option model=m --no-verify \ | |
| --compact --rows shards/<i>.jsonl.gz --out shards/<i> # i = 0..63, run in parallel; 16 shards per server | |
| cat shards/*/results.jsonl > results.jsonl | |
| decision-index score --suite-dir suite-0.3 --results results.jsonl --engine http --out score | |
| ``` | |
| ## Hardware and latency | |
| - Four NVIDIA H20 (96 GB each), driver 590.48.01, one server per GPU. | |
| - Gen3 has the same architecture and size as Gen1 and Gen2, so its per-request compute is the same. | |
| - Latency was not measured on the board's reference hardware (1x RTX PRO 6000, latency-v1). | |
| ## One-command Decision Index run | |
| `knowline_engine.py` (in this repository) is a Decision Index engine. It runs the same front end in process, so there | |
| is no server to start by hand: | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests # plus the decision-index kit | |
| git clone https://huggingface.co/PelaAI/KnowLine-4B-Gen3 && cd KnowLine-4B-Gen3 # puts both files on the path | |
| python -m decision_index pipeline --engine knowline_engine:KnowLine \ | |
| --option model=PelaAI/KnowLine-4B-Gen3 --option revision=<commit> --out runs/KnowLine-4B-Gen3 | |
| ``` | |
| - **Default backend (the setting of our runs):** the engine starts SGLang with the flags above (FP8 at load) on a free | |
| local port and stops it when the run ends. Use `--option gpu=<n>` to pick a GPU, and | |
| `--option sglang_python=<python>` if SGLang lives in another environment. | |
| - **`--option backend=hf`:** transformers only, bf16, no server; slower, and not the setting of our runs. | |
| - **Limits:** up to 64 questions per request and 255 options per question. Larger requests are reported as | |
| unsupported; nothing is truncated. | |
| - **Checked on Gen2** (same architecture, front end and engine file): on 300 random Decision Index 0.3 rows, the | |
| default backend gave the same answer as our published Gen2 run on 706 of 709 questions. The 3 differences are | |
| near-ties (top two options within 0.06), from FP8 numerical noise. Requests took 30.7 ms median, one at a time on one | |
| H20. | |
| ## Changes | |
| - 2026-10-08 (also in Gen1 and Gen2): `knowline_server.py` accepts chats whose roles the chat template rejects (for | |
| example `customer` / `agent`): such a state is rendered as one user message, like any other structured state, instead | |
| of the request failing. Any other unexpected error returns HTTP 500 instead of closing the connection. Requests that | |
| worked before give the same answers. | |