Instructions to use PelaAI/KnowLine-4B-Gen1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen1")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen1") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download INFERENCE.md from PelaAI/KnowLine-4B-Gen1: direct link, hf CLI and curl.
- Browser
- Download file 6.94 kB
-
https://huggingface.co/PelaAI/KnowLine-4B-Gen1/resolve/main/INFERENCE.md
- Command line
-
hf download hf://PelaAI/KnowLine-4B-Gen1/INFERENCE.md
-
curl -L -o INFERENCE.md https://huggingface.co/PelaAI/KnowLine-4B-Gen1/resolve/main/INFERENCE.md
KnowLine-4B-Gen1: serving and Decision Index reproduction
These settings reproduce our self-run Decision Index evaluation of these weights: run openjev-4b-e-best, 0.2.1 scored
2026-10-07 03:22 CST, and 0.3 public with the 2,638 rebuilt GSM8K rows added.
Weights
- This repository holds the merged bf16 weights: Qwen3.5-4B plus our LoRA, step 5,650 of training run "mix E".
- The base model's vision tower and MTP head are unchanged. The architecture is
Qwen3_5ForConditionalGeneration, with the base model'sconfig.json. SHA256SUMSlists every file.
Software
| component | version |
|---|---|
| SGLang | 0.5.21 (torch 2.13.0, CUDA 13.0, flashinfer-python 0.6.18) |
| transformers | 5.12.1 |
| front end | knowline_server.py (this repository; needs only transformers and requests) |
| decision-index kit | 0.2.1 at commit 87d4650b42b377c0291a89c1f1a879f9b31082bf; 0.3 at commit 62d2f51de34a2de64906345b6bc3e98e27ff55c7 |
About the front end
knowline_server.py is the /v1/systemone front end of our run, packaged as one file. It contains:
- rendering: the chat template with thinking off; the state is followed by one user turn with the instruction, the
question and all options labelled A, B, ...; the assistant turn opens with
Answer:; - label-token scoring: one prefill per question, then a softmax over the option labels only;
- the HTTP server.
We checked it against the exact front end that produced our Decision Index run:
- prompts, answer keys and label token ids were identical on 7,500 Decision Index rows and training requests;
- through a live SGLang engine, choices agreed on all 953 questions of 200 Decision Index rows. The largest probability difference, 0.02, is the same as between two identical calls to one engine (FP8 batching noise).
Serve
serve_knowline.sh starts both processes.
- SGLang engine.
- FP8 is on at serving: weights are stored in bf16 and SGLang quantizes them on load.
--mem-fraction-static 0.72only sizes SGLang's KV cache. Any value works and scores do not depend on it.- SGLang treats the model as multimodal by itself. The Decision Index run sent text only.
/v1/systemonefront end:knowline_server.py,chatstyle, temperature 1, 16 scoring threads, no per-type calibration file.
pip install "sglang==0.5.21" "transformers==5.12.1" requests # torch / CUDA per SGLang's install docs
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen1 0 8080 # GPU 0; SGLang on :9080, /v1/systemone on :8080
curl -s http://127.0.0.1:8080/health
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "m",
"state": "Customer: my order arrived broken, I want my money back.",
"questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
"tone": {"type": "choice", "instructions": "Customer tone?",
"criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
Equivalent manual launch, with the exact flags of our run:
CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server --model-path PelaAI/KnowLine-4B-Gen1 --served-model-name m --tp 1 \
--quantization fp8 --mem-fraction-static 0.72 --mamba-radix-cache-strategy extra_buffer --enable-fp32-lm-head --port 9080 &
python knowline_server.py --model PelaAI/KnowLine-4B-Gen1 --backend sglang --url http://127.0.0.1:9080 \
--served-model-name m --temperature 1 --workers 16 --port 8080
Without SGLang (CPU or a single GPU through transformers; slower, no prefix cache, bf16 rather than FP8):
pip install torch "transformers==5.12.1" requests
python knowline_server.py --model PelaAI/KnowLine-4B-Gen1 --backend hf --port 8080
Decision Index run
- Suite rows:
suite-0.2/selected-rows.jsonl.gzandadded-rows.jsonl.gz, 155,390 rows in total. - The rows are split round-robin into 16 shards. Each shard is a resumable
decision-index runagainst the front end:
decision-index run --engine http --option base_url=http://127.0.0.1:8080 --option model=m --no-verify --compact \
--rows shards/<i>.jsonl.gz --out shards/<i> # i = 0..15, run in parallel
cat shards/*/results.jsonl > results.jsonl
decision-index score --suite-dir suite-0.2 --results results.jsonl --engine http --out score
- Servers: the run started on one server (01:56 CST). From 01:59 CST it used two identical servers on two GPUs, shards 0-7 and 8-15; that changes speed, not answers.
- 0.3 public score: the same 0.2.1 run plus the 2,638 rebuilt GSM8K rows (kit 0.3,
suite rebuild), scored with--edition 0.3. The run directory is delivered with the submission.
Hardware and latency
- NVIDIA H20 (96 GB), driver 590.48.01, one server per GPU.
- Latency was not measured on the board's reference hardware (1x RTX PRO 6000, latency-v1).
One-command Decision Index run
knowline_engine.py (in this repository) is a Decision Index engine. It runs the same front end in process, so there
is no server to start by hand:
pip install "sglang==0.5.21" "transformers==5.12.1" requests # plus the decision-index kit
git clone https://huggingface.co/PelaAI/KnowLine-4B-Gen1 && cd KnowLine-4B-Gen1 # puts both files on the path
python -m decision_index pipeline --engine knowline_engine:KnowLine \
--option model=PelaAI/KnowLine-4B-Gen1 --option revision=<commit> --out runs/KnowLine-4B-Gen1
- Default backend (the setting of our runs): the engine starts SGLang with the flags above (FP8 at load) on a free
local port and stops it when the run ends. Use
--option gpu=<n>to pick a GPU, and--option sglang_python=<python>if SGLang lives in another environment. --option backend=hf: transformers only, bf16, no server; slower, and not the setting of our runs.- Limits: up to 64 questions per request and 255 options per question. Larger requests are reported as unsupported; nothing is truncated.
- Checked on Gen2: on 300 random Decision Index 0.3 rows, the default backend gave the same answer as our published Gen2 run on 706 of 709 questions. The 3 differences are near-ties (top two options within 0.06), from FP8 numerical noise. Requests took 30.7 ms median, one at a time on one H20.
Changes
- 2026-10-08:
knowline_server.pynow accepts chats whose roles the chat template rejects (for examplecustomer/agent): such a state is rendered as one user message, like any other structured state, instead of the request failing. Any other unexpected error now returns HTTP 500 instead of closing the connection. Requests that worked before give the same answers; our Decision Index runs had none that failed. - 2026-10-08: added
knowline_engine.py(one-command Decision Index runs).--backend hfofknowline_server.pyno longer needsaccelerate: without it, the model is loaded and moved to one device.