File size: 8,763 Bytes
6fad739
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e088461
 
 
3e73f0e
6fad739
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47a81a0
e088461
3e73f0e
6fad739
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3e73f0e
 
 
6fad739
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
---
license: apache-2.0
base_model:
  - Qwen/Qwen3.5-4B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-classification
language:
  - en
  - zh
tags:
  - decision-model
  - system-one
  - decision-index
  - lora-merged
---

# KnowLine-4B-Gen1

**A 4B System One decision model from PelaAI, trained on a single all-in-one machine by an agent-driven data loop.**

[Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0

> **Newer versions:** [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) (Decision Index 0.3 public score 63.11) and
> [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) (62.54) are out. Both are at or above this model on our held-out, C-Eval, Open-Jev and prompt-injection
> evaluations.

Highlights:

- **Decision Index 0.3, public suite: 60.47** (self-run). For comparison: Jev 1.13 scores 57.96 (board), and the best
  ≤5B model on the board has a public score of 50.82; see [Comparison](#comparison).
- **An AI agent ran the whole loop.** In each round it:
  - found the areas where the model was weak;
  - proposed a targeted data group;
  - built and decontaminated the data;
  - trained and evaluated on it;
  - kept or rejected the change based on the evidence.
- **KOF '98 harness: 17 wins, 0 losses, 1 draw against Jev 1.13** in a single-bout mirror round robin. For the play
  style, see [Game harness](#game-harness-kof-98).

KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev.

## What it does

You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model
answers every question with a probability distribution over its options:

- one forward pass per question, with no generated text;
- the output is the probability of each option's label token;
- existing Jev clients only need a new base URL.

| | |
|---|---|
| Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) |
| Training | LoRA SFT (rank 32, alpha 64) on the language model, one epoch over about 713k rows. This release is step 5,650 of 5,655, the lowest validation loss. |
| Release format | LoRA merged into the base; bf16 weights. The vision tower and MTP head are unchanged. |
| Languages | English, Simplified Chinese, Traditional Chinese |

## Quickstart

```bash
pip install "sglang==0.5.21" "transformers==5.12.1" requests
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen1 0 8080     # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "m",
  "state": "Customer: my order arrived broken, I want my money back.",
  "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
                "tone": {"type": "choice", "instructions": "Customer tone?",
                         "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
```

- **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen1 --backend hf --port 8080` uses
  transformers only. It is slower and runs bf16.
- **Front end:** `knowline_server.py` is a single file and needs only transformers and requests.
- **Full settings:** the exact settings of our Decision Index runs are in [INFERENCE.md](INFERENCE.md).

## Comparison

Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill,
0.3 new-domain); ours is not available yet.

| model | size | DI 0.3 public | DI 0.3 full | source |
|---|---|---|---|---|
| **KnowLine-4B-Gen1** | 4B | **60.47** | not yet scored | self-run, official kit |
| Clef | 27B | 61.71 | 53.08 | board |
| Jev 1.13 | (API) | 57.96 | 60.11 | board |
| RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported |
| ezjev 4B s2 | 4B | 50.82 | 46.95 | board |
| jiwo 4B | 4B | 45.76 | 42.86 | board |
| Nox 4B | 4B | 44.21 | 44.95 | board |

Our public-suite scores by area:

| area | KnowLine-4B-Gen1 (0.3 public) | Jev 1.13 (0.2.1) |
|---|---|---|
| Knowledge & reasoning | 37.7 | 51.4 |
| Language | 60.8 | 62.0 |
| Retrieval & routing | 70.9 | 55.4 |
| Tools & agents | 86.8 | 75.1 |
| Arts & taste | 49.5 | 37.7 |

## Releases

The model is trained in a self-evolving loop, and new versions will follow.

| model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes |
|---|---|---|---|---|---|
| [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release |
| [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release |
| [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | this model at 0.5 + mix F steps 4,500 and 5,000 at 0.25 each, weight average |
| KnowLine-4B-Gen1 (this model) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release (internal run "mix E", step 5,650) |

## Evaluation

All results are self-run and not verified by a third party.

### Held-out and Chinese evaluations

| suite | accuracy |
|---|---|
| In-house evaluation set, English | 68.7 |
| In-house evaluation set, Simplified Chinese | 69.9 |
| In-house evaluation set, Traditional Chinese | 64.1 |
| C-Eval (4 categories, macro) | 77.9 |
| Open-Jev 1.1 test / OOD | 87.4 / 86.6 |

### Web operation (Mind2Web official test splits, evaluation only)

| split | element selection | operation (balanced) |
|---|---|---|
| test_task (websites seen in training, new tasks) | 93.1 | 96.6 |
| test_website (new websites) | 90.7 | 96.5 |
| test_domain (new domains) | 91.5 | 98.5 |

- So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree.
- The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier
  than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard.

### Game harness (KOF '98)

- **Setup:** single-bout character-mirror round robin, 11 players, 18 games per pair (990 games), argmax actions.
- **Result:** win rate 0.883 [0.83, 0.92], Elo 1910, tied for first of 11 with another internal checkpoint.
- **Head-to-head:** 17-0-1 against Jev 1.13, 15-3 against StartLux-Decision-4B, 18-0 against Clef-Flash.
- **Caveat:** the policy relies heavily on one move, a 623C anti-air uppercut used about 55% of the time. This is a
  measured result in this harness, not general fighting-game skill.

## Disclosures

- **Decision Index training data:** about 25% of the data is in Decision Index format, that is, synthetic items written
  in the benchmarks' request formats. No Decision Index test item is included.
- **Decontamination:**
  - Every component text (state, instructions, option texts) was checked against every Decision Index row and all of
    our evaluation sets.
  - Rows that are identical, or share any run of 50 consecutive characters, were removed.
- **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation
  states.
- **No model outputs as labels:** no output of Jev or any other decision model was used as a training label.
- **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but
  not all were checked by a human.

## Limitations

- **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 37.7 (0.3); MMLU-Pro, GPQA and HLE are
  close to the base model.
- **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 29.5.
- **Prompt injection:** an instruction planted in the state changes the answer about 9% of the time on our injection
  set. Keep untrusted text clearly delimited.
- **Calibration:** ECE is about 0.05-0.06 on choice and yes/no questions and about 0.12 on score questions. For score
  questions, fit a temperature on your own data. On the Decision Index 0.3 public suite, computed the board's way, ECE
  is about 0.08-0.09 and 3.3% of answers are wrong at ≥95% confidence (board median 1.3%): the model is overconfident
  overall.
- **Private tests:** part of our public-suite advantage comes from adapting to the question formats, so the 0.3 private
  tests may score lower.

## Citation

```bibtex
@misc{knowline4bgen1,
  title  = {KnowLine-4B-Gen1: a 4B decision model},
  author = {PelaAI},
  year   = {2026},
  url    = {https://huggingface.co/PelaAI/KnowLine-4B-Gen1}
}
```

## License

- **Weights:** Apache-2.0, the same as the base model.
- **Code:** `knowline_server.py` is MIT (see the file header).