KnowLine-4B-Gen3 / INFERENCE.md
PEScn's picture
KnowLine-4B-Gen3: 0.5 Gen2 + 0.5 run g3 step 8000 (bf16)
0022dbc verified
|
Raw History Blame Contribute Delete
7.25 kB

KnowLine-4B-Gen3: serving and Decision Index reproduction

These settings reproduce our self-run Decision Index 0.3 evaluation of these weights: the full 0.3 public suite sent fresh (no rows carried over from another run), scored 2026-10-08 10:42 CST.

Weights

  • This repository holds bf16 weights made by averaging merged models tensor by tensor (each step computed in fp32 and stored in bf16):
    • KnowLine-4B-Gen2, 0.5;
    • a model from the next training round (Qwen3.5-4B plus a LoRA r32, merged), 0.5.
  • 248 tensors are averaged. The other 490 are identical in both models and copied unchanged, including the base model's vision tower and MTP head.
  • The architecture is Qwen3_5ForConditionalGeneration. config.json, the tokenizer files and the chat template are byte-identical to Gen1 and Gen2.
  • SHA256SUMS lists every file.

Software

component version
SGLang 0.5.21 (torch 2.13.0, CUDA 13.0, flashinfer-python 0.6.18)
transformers 5.12.1
front end knowline_server.py (this repository; needs only transformers and requests)
decision-index kit 0.3 at commit 62d2f51de34a2de64906345b6bc3e98e27ff55c7

About the front end

knowline_server.py is the /v1/systemone front end of our runs, packaged as one file. It is the same file as in Gen2; only the model name in its docstring changed. It contains:

  • rendering: the chat template with thinking off; the state is followed by one user turn with the instruction, the question and all options labelled A, B, ...; the assistant turn opens with Answer:;
  • label-token scoring: one prefill per question, then a softmax over the option labels only;
  • the HTTP server.

The Gen3 run was served by the same in-house front end as the Gen1 and Gen2 runs. Before the Gen3 run it received the same 2026-10-08 fix as knowline_server.py (see Changes), which gives the same answers on every request that worked before. knowline_server.py was checked against the in-house front end on Gen1:

  • prompts, answer keys and label token ids were identical on 7,500 Decision Index rows and training requests;
  • through a live SGLang engine, choices agreed on all 953 questions of 200 Decision Index rows.

Serve

serve_knowline.sh starts both processes.

  1. SGLang engine.
    • FP8 is on at serving: weights are stored in bf16 and SGLang quantizes them on load.
    • --mem-fraction-static 0.72 only sizes SGLang's KV cache. Any value works and scores do not depend on it.
    • SGLang treats the model as multimodal by itself. The Decision Index run sent text only.
  2. /v1/systemone front end: knowline_server.py, chat style, temperature 1, 16 scoring threads, no per-type calibration file.
pip install "sglang==0.5.21" "transformers==5.12.1" requests   # torch / CUDA per SGLang's install docs
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen3 0 8080           # GPU 0; SGLang on :9080, /v1/systemone on :8080
curl -s http://127.0.0.1:8080/health
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "m",
  "state": "Customer: my order arrived broken, I want my money back.",
  "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
                "tone": {"type": "choice", "instructions": "Customer tone?",
                         "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'

Equivalent manual launch, with the exact flags of our run:

CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server --model-path PelaAI/KnowLine-4B-Gen3 --served-model-name m --tp 1 \
  --quantization fp8 --mem-fraction-static 0.72 --mamba-radix-cache-strategy extra_buffer --enable-fp32-lm-head --port 9080 &
python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend sglang --url http://127.0.0.1:9080 \
  --served-model-name m --temperature 1 --workers 16 --port 8080

Without SGLang (CPU or a single GPU through transformers; slower, no prefix cache, bf16 rather than FP8):

pip install torch "transformers==5.12.1" requests
python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend hf --port 8080

Decision Index run

  • Suite rows: the 140,620 rows that decision-index run --edition 0.3 sends: selected-rows.jsonl.gz, added-rows.jsonl.gz and gsm8k-rows.jsonl.gz of suite-0.3, with ToolRet, BRIGHT and ACOS cut to their scoring subsets.
  • Four identical servers ran on four GPUs, one each, with the settings above. The rows are split round-robin into 64 shards, 16 per server. Each shard is a resumable decision-index run against its server:
decision-index run --edition 0.3 --engine http --option base_url=http://127.0.0.1:<port> --option model=m --no-verify \
  --compact --rows shards/<i>.jsonl.gz --out shards/<i>          # i = 0..63, run in parallel; 16 shards per server
cat shards/*/results.jsonl > results.jsonl
decision-index score --suite-dir suite-0.3 --results results.jsonl --engine http --out score

Hardware and latency

  • Four NVIDIA H20 (96 GB each), driver 590.48.01, one server per GPU.
  • Gen3 has the same architecture and size as Gen1 and Gen2, so its per-request compute is the same.
  • Latency was not measured on the board's reference hardware (1x RTX PRO 6000, latency-v1).

One-command Decision Index run

knowline_engine.py (in this repository) is a Decision Index engine. It runs the same front end in process, so there is no server to start by hand:

pip install "sglang==0.5.21" "transformers==5.12.1" requests        # plus the decision-index kit
git clone https://huggingface.co/PelaAI/KnowLine-4B-Gen3 && cd KnowLine-4B-Gen3   # puts both files on the path
python -m decision_index pipeline --engine knowline_engine:KnowLine \
    --option model=PelaAI/KnowLine-4B-Gen3 --option revision=<commit> --out runs/KnowLine-4B-Gen3
  • Default backend (the setting of our runs): the engine starts SGLang with the flags above (FP8 at load) on a free local port and stops it when the run ends. Use --option gpu=<n> to pick a GPU, and --option sglang_python=<python> if SGLang lives in another environment.
  • --option backend=hf: transformers only, bf16, no server; slower, and not the setting of our runs.
  • Limits: up to 64 questions per request and 255 options per question. Larger requests are reported as unsupported; nothing is truncated.
  • Checked on Gen2 (same architecture, front end and engine file): on 300 random Decision Index 0.3 rows, the default backend gave the same answer as our published Gen2 run on 706 of 709 questions. The 3 differences are near-ties (top two options within 0.06), from FP8 numerical noise. Requests took 30.7 ms median, one at a time on one H20.

Changes

  • 2026-10-08 (also in Gen1 and Gen2): knowline_server.py accepts chats whose roles the chat template rejects (for example customer / agent): such a state is rendered as one user message, like any other structured state, instead of the request failing. Any other unexpected error returns HTTP 500 instead of closing the connection. Requests that worked before give the same answers.