--- license: apache-2.0 base_model: - Qwen/Qwen3.5-4B base_model_relation: finetune library_name: transformers pipeline_tag: text-classification language: - en - zh tags: - decision-model - system-one - decision-index - lora-merged --- # KnowLine-4B-Gen1 **A 4B System One decision model from PelaAI, trained on a single all-in-one machine by an agent-driven data loop.** [Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0 > **Newer versions:** [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) (Decision Index 0.3 public score 63.11) and > [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) (62.54) are out. Both are at or above this model on our held-out, C-Eval, Open-Jev and prompt-injection > evaluations. Highlights: - **Decision Index 0.3, public suite: 60.47** (self-run). For comparison: Jev 1.13 scores 57.96 (board), and the best ≤5B model on the board has a public score of 50.82; see [Comparison](#comparison). - **An AI agent ran the whole loop.** In each round it: - found the areas where the model was weak; - proposed a targeted data group; - built and decontaminated the data; - trained and evaluated on it; - kept or rejected the change based on the evidence. - **KOF '98 harness: 17 wins, 0 losses, 1 draw against Jev 1.13** in a single-bout mirror round robin. For the play style, see [Game harness](#game-harness-kof-98). KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev. ## What it does You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model answers every question with a probability distribution over its options: - one forward pass per question, with no generated text; - the output is the probability of each option's label token; - existing Jev clients only need a new base URL. | | | |---|---| | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) | | Training | LoRA SFT (rank 32, alpha 64) on the language model, one epoch over about 713k rows. This release is step 5,650 of 5,655, the lowest validation loss. | | Release format | LoRA merged into the base; bf16 weights. The vision tower and MTP head are unchanged. | | Languages | English, Simplified Chinese, Traditional Chinese | ## Quickstart ```bash pip install "sglang==0.5.21" "transformers==5.12.1" requests bash serve_knowline.sh PelaAI/KnowLine-4B-Gen1 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080 curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{ "model": "m", "state": "Customer: my order arrived broken, I want my money back.", "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"}, "tone": {"type": "choice", "instructions": "Customer tone?", "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}' ``` - **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen1 --backend hf --port 8080` uses transformers only. It is slower and runs bf16. - **Front end:** `knowline_server.py` is a single file and needs only transformers and requests. - **Full settings:** the exact settings of our Decision Index runs are in [INFERENCE.md](INFERENCE.md). ## Comparison Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill, 0.3 new-domain); ours is not available yet. | model | size | DI 0.3 public | DI 0.3 full | source | |---|---|---|---|---| | **KnowLine-4B-Gen1** | 4B | **60.47** | not yet scored | self-run, official kit | | Clef | 27B | 61.71 | 53.08 | board | | Jev 1.13 | (API) | 57.96 | 60.11 | board | | RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported | | ezjev 4B s2 | 4B | 50.82 | 46.95 | board | | jiwo 4B | 4B | 45.76 | 42.86 | board | | Nox 4B | 4B | 44.21 | 44.95 | board | Our public-suite scores by area: | area | KnowLine-4B-Gen1 (0.3 public) | Jev 1.13 (0.2.1) | |---|---|---| | Knowledge & reasoning | 37.7 | 51.4 | | Language | 60.8 | 62.0 | | Retrieval & routing | 70.9 | 55.4 | | Tools & agents | 86.8 | 75.1 | | Arts & taste | 49.5 | 37.7 | ## Releases The model is trained in a self-evolving loop, and new versions will follow. | model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes | |---|---|---|---|---|---| | [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release | | [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release | | [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | this model at 0.5 + mix F steps 4,500 and 5,000 at 0.25 each, weight average | | KnowLine-4B-Gen1 (this model) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release (internal run "mix E", step 5,650) | ## Evaluation All results are self-run and not verified by a third party. ### Held-out and Chinese evaluations | suite | accuracy | |---|---| | In-house evaluation set, English | 68.7 | | In-house evaluation set, Simplified Chinese | 69.9 | | In-house evaluation set, Traditional Chinese | 64.1 | | C-Eval (4 categories, macro) | 77.9 | | Open-Jev 1.1 test / OOD | 87.4 / 86.6 | ### Web operation (Mind2Web official test splits, evaluation only) | split | element selection | operation (balanced) | |---|---|---| | test_task (websites seen in training, new tasks) | 93.1 | 96.6 | | test_website (new websites) | 90.7 | 96.5 | | test_domain (new domains) | 91.5 | 98.5 | - So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree. - The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard. ### Game harness (KOF '98) - **Setup:** single-bout character-mirror round robin, 11 players, 18 games per pair (990 games), argmax actions. - **Result:** win rate 0.883 [0.83, 0.92], Elo 1910, tied for first of 11 with another internal checkpoint. - **Head-to-head:** 17-0-1 against Jev 1.13, 15-3 against StartLux-Decision-4B, 18-0 against Clef-Flash. - **Caveat:** the policy relies heavily on one move, a 623C anti-air uppercut used about 55% of the time. This is a measured result in this harness, not general fighting-game skill. ## Disclosures - **Decision Index training data:** about 25% of the data is in Decision Index format, that is, synthetic items written in the benchmarks' request formats. No Decision Index test item is included. - **Decontamination:** - Every component text (state, instructions, option texts) was checked against every Decision Index row and all of our evaluation sets. - Rows that are identical, or share any run of 50 consecutive characters, were removed. - **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation states. - **No model outputs as labels:** no output of Jev or any other decision model was used as a training label. - **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but not all were checked by a human. ## Limitations - **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 37.7 (0.3); MMLU-Pro, GPQA and HLE are close to the base model. - **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 29.5. - **Prompt injection:** an instruction planted in the state changes the answer about 9% of the time on our injection set. Keep untrusted text clearly delimited. - **Calibration:** ECE is about 0.05-0.06 on choice and yes/no questions and about 0.12 on score questions. For score questions, fit a temperature on your own data. On the Decision Index 0.3 public suite, computed the board's way, ECE is about 0.08-0.09 and 3.3% of answers are wrong at ≥95% confidence (board median 1.3%): the model is overconfident overall. - **Private tests:** part of our public-suite advantage comes from adapting to the question formats, so the 0.3 private tests may score lower. ## Citation ```bibtex @misc{knowline4bgen1, title = {KnowLine-4B-Gen1: a 4B decision model}, author = {PelaAI}, year = {2026}, url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen1} } ``` ## License - **Weights:** Apache-2.0, the same as the base model. - **Code:** `knowline_server.py` is MIT (see the file header).