Text Classification
Transformers
Safetensors
English
Chinese
qwen3_5
image-text-to-text
decision-model
system-one
decision-index
lora-merged
model-soup
Instructions to use PelaAI/KnowLine-4B-Gen2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen2")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen2") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download INFERENCE.md from PelaAI/KnowLine-4B-Gen2: direct link, hf CLI and curl.
- Browser
- Download file 7.35 kB
-
https://huggingface.co/PelaAI/KnowLine-4B-Gen2/resolve/main/INFERENCE.md
- Command line
-
hf download hf://PelaAI/KnowLine-4B-Gen2/INFERENCE.md
-
curl -L -o INFERENCE.md https://huggingface.co/PelaAI/KnowLine-4B-Gen2/resolve/main/INFERENCE.md
7.35 kB
| # KnowLine-4B-Gen2: serving and Decision Index reproduction | |
| These settings reproduce our self-run Decision Index 0.3 evaluation of these weights: run | |
| `openjev-4b-soup-e5650-f4500f5000-v03`, the full 0.3 public suite sent fresh (no rows carried over from another run), | |
| scored 2026-10-07 22:38 CST. | |
| ## Weights | |
| - This repository holds bf16 weights made by averaging merged models tensor by tensor (each step computed in fp32 and | |
| stored in bf16), in two steps: | |
| 1. steps 4,500 and 5,000 of training run "mix F" (Qwen3.5-4B plus a LoRA r32, merged), 0.5 each; | |
| 2. that average and [KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1) (Qwen3.5-4B plus our LoRA, | |
| step 5,650 of training run "mix E"), 0.5 each. | |
| - So Gen2 = 0.5 × Gen1 + 0.25 × mix F step 4,500 + 0.25 × mix F step 5,000. | |
| - In each step 248 tensors are averaged. The other 490 are identical in all models and copied unchanged, including the | |
| base model's vision tower and MTP head. | |
| - The architecture is `Qwen3_5ForConditionalGeneration`. `config.json`, the tokenizer files and the chat template are | |
| byte-identical to Gen1. | |
| - `SHA256SUMS` lists every file. | |
| ## Software | |
| | component | version | | |
| |---|---| | |
| | SGLang | 0.5.21 (torch 2.13.0, CUDA 13.0, flashinfer-python 0.6.18) | | |
| | transformers | 5.12.1 | | |
| | front end | `knowline_server.py` (this repository; needs only transformers and requests) | | |
| | decision-index kit | 0.3 at commit 62d2f51de34a2de64906345b6bc3e98e27ff55c7 | | |
| ### About the front end | |
| `knowline_server.py` is the `/v1/systemone` front end of our runs, packaged as one file. It is the same file as in | |
| Gen1; only the model name in its docstring changed. It contains: | |
| - rendering: the chat template with thinking off; the state is followed by one user turn with the instruction, the | |
| question and all options labelled A, B, ...; the assistant turn opens with `Answer:`; | |
| - label-token scoring: one prefill per question, then a softmax over the option labels only; | |
| - the HTTP server. | |
| The Gen2 run was served by the same in-house front end as the Gen1 run. `knowline_server.py` was checked against it on | |
| Gen1: | |
| - prompts, answer keys and label token ids were identical on 7,500 Decision Index rows and training requests; | |
| - through a live SGLang engine, choices agreed on all 953 questions of 200 Decision Index rows. | |
| ## Serve | |
| `serve_knowline.sh` starts both processes. | |
| 1. **SGLang engine.** | |
| - FP8 is on at serving: weights are stored in bf16 and SGLang quantizes them on load. | |
| - `--mem-fraction-static 0.72` only sizes SGLang's KV cache. Any value works and scores do not depend on it. | |
| - SGLang treats the model as multimodal by itself. The Decision Index run sent text only. | |
| 2. **`/v1/systemone` front end:** `knowline_server.py`, `chat` style, temperature 1, 16 scoring threads, no per-type | |
| calibration file. | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests # torch / CUDA per SGLang's install docs | |
| bash serve_knowline.sh PelaAI/KnowLine-4B-Gen2 0 8080 # GPU 0; SGLang on :9080, /v1/systemone on :8080 | |
| curl -s http://127.0.0.1:8080/health | |
| curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{ | |
| "model": "m", | |
| "state": "Customer: my order arrived broken, I want my money back.", | |
| "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"}, | |
| "tone": {"type": "choice", "instructions": "Customer tone?", | |
| "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}' | |
| ``` | |
| Equivalent manual launch, with the exact flags of our run: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server --model-path PelaAI/KnowLine-4B-Gen2 --served-model-name m --tp 1 \ | |
| --quantization fp8 --mem-fraction-static 0.72 --mamba-radix-cache-strategy extra_buffer --enable-fp32-lm-head --port 9080 & | |
| python knowline_server.py --model PelaAI/KnowLine-4B-Gen2 --backend sglang --url http://127.0.0.1:9080 \ | |
| --served-model-name m --temperature 1 --workers 16 --port 8080 | |
| ``` | |
| Without SGLang (CPU or a single GPU through transformers; slower, no prefix cache, bf16 rather than FP8): | |
| ```bash | |
| pip install torch "transformers==5.12.1" requests | |
| python knowline_server.py --model PelaAI/KnowLine-4B-Gen2 --backend hf --port 8080 | |
| ``` | |
| ## Decision Index run | |
| - Suite rows: the 140,620 rows that `decision-index run --edition 0.3` sends: `selected-rows.jsonl.gz`, | |
| `added-rows.jsonl.gz` and `gsm8k-rows.jsonl.gz` of `suite-0.3`, with ToolRet, BRIGHT and ACOS cut to their scoring | |
| subsets. | |
| - The rows are split round-robin into 16 shards. Each shard is a resumable `decision-index run` against one server: | |
| ```bash | |
| decision-index run --edition 0.3 --engine http --option base_url=http://127.0.0.1:8080 --option model=m --no-verify \ | |
| --compact --rows shards/<i>.jsonl.gz --out shards/<i> # i = 0..15, run in parallel | |
| cat shards/*/results.jsonl > results.jsonl | |
| decision-index score --suite-dir suite-0.3 --results results.jsonl --engine http --out score | |
| ``` | |
| ## Hardware and latency | |
| - One NVIDIA H20 (96 GB), driver 590.48.01, one server for all 16 shards. | |
| - Gen2 has the same architecture and size as Gen1, so its per-request compute is the same. | |
| - Latency was not measured on the board's reference hardware (1x RTX PRO 6000, latency-v1). | |
| ## One-command Decision Index run | |
| `knowline_engine.py` (in this repository) is a Decision Index engine. It runs the same front end in process, so there | |
| is no server to start by hand: | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests # plus the decision-index kit | |
| git clone https://huggingface.co/PelaAI/KnowLine-4B-Gen2 && cd KnowLine-4B-Gen2 # puts both files on the path | |
| python -m decision_index pipeline --engine knowline_engine:KnowLine \ | |
| --option model=PelaAI/KnowLine-4B-Gen2 --option revision=<commit> --out runs/KnowLine-4B-Gen2 | |
| ``` | |
| - **Default backend (the setting of our runs):** the engine starts SGLang with the flags above (FP8 at load) on a free | |
| local port and stops it when the run ends. Use `--option gpu=<n>` to pick a GPU, and | |
| `--option sglang_python=<python>` if SGLang lives in another environment. | |
| - **`--option backend=hf`:** transformers only, bf16, no server; slower, and not the setting of our runs. | |
| - **Limits:** up to 64 questions per request and 255 options per question. Larger requests are reported as | |
| unsupported; nothing is truncated. | |
| - **Checked:** on 300 random Decision Index 0.3 rows, the default backend gave the same answer as our published | |
| Gen2 run on 706 of 709 questions. The 3 differences are near-ties (top two options within 0.06), from FP8 numerical | |
| noise. Requests took 30.7 ms median, one at a time on one H20. | |
| ## Changes | |
| - 2026-10-08: `knowline_server.py` now accepts chats whose roles the chat template rejects (for example `customer` / | |
| `agent`): such a state is rendered as one user message, like any other structured state, instead of the request | |
| failing. Any other unexpected error now returns HTTP 500 instead of closing the connection. Requests that worked | |
| before give the same answers; our Decision Index runs had none that failed. | |
| - 2026-10-08: added `knowline_engine.py` (one-command Decision Index runs). `--backend hf` of `knowline_server.py` no | |
| longer needs `accelerate`: without it, the model is loaded and moved to one device. | |