File size: 14,103 Bytes
c7ca0de
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2ac7b70
 
 
c7ca0de
 
 
 
 
 
 
 
 
b0e507b
 
c7ca0de
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ede2694
2ac7b70
230724f
c7ca0de
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b0e507b
c7ca0de
b0e507b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c7ca0de
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
---
license: apache-2.0
base_model:
  - Qwen/Qwen3.5-4B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-classification
language:
  - en
  - zh
tags:
  - decision-model
  - system-one
  - decision-index
  - lora-merged
  - model-soup
---

# KnowLine-4B-Gen2

**The second release of PelaAI's 4B System One decision model: a weight average of KnowLine-4B-Gen1 and the model from
the next training round.**

[Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0 · previous version:
[KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1)

> **Newer version:** [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) is out, with a Decision Index 0.3 public score of
> 63.11.

Highlights:

- **Decision Index 0.3, public suite: 62.54** (self-run), 2.07 above Gen1 (60.47). For comparison: Jev 1.13 scores
  57.96 (board), and the best ≤5B model on the board has a public score of 50.82; see [Comparison](#comparison).
- **At or above Gen1 on all of our own evaluations:** in-house held-out set (en / zh-Hans / zh-Hant) 69.7 / 71.2 / 64.8
  (Gen1: 68.7 / 69.9 / 64.1), C-Eval 78.6 (77.9). Instructions planted in the state change the answer 4.0% of the
  time, down from 8.9%.
- **Better calibrated:** computed the board's way, ECE is about 0.06-0.07 (Gen1 about 0.08-0.09, board median 0.084),
  and the Brier score is about 0.31-0.33, roughly 3rd of 113 models on the board; see [Calibration](#calibration).
- **KOF '98 harness: 14-3-1 against Jev 1.13 and 13-5 against Gen1** in single-bout mirror matches; see
  [Game harness](#game-harness-kof-98).
- **An AI agent ran the whole loop.** In each round it:
  - found the areas where the model was weak;
  - proposed a targeted data group;
  - built and decontaminated the data;
  - trained and evaluated on it;
  - kept or rejected the change based on the evidence.

  For this round, see [What changed in Gen2](#what-changed-in-gen2).

KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev.

## What it does

You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model
answers every question with a probability distribution over its options:

- one forward pass per question, with no generated text;
- the output is the probability of each option's label token;
- existing Jev clients only need a new base URL; Gen1 users only change the model name, since the interface and
  serving settings are the same.

| | |
|---|---|
| Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) |
| Training | Two LoRA SFT runs (rank 32, alpha 64, language model only), merged, then averaged tensor by tensor: Gen1 (run "mix E", step 5,650) at 0.5, and steps 4,500 and 5,000 of run "mix F" at 0.25 each. |
| Release format | bf16 weights. The vision tower and MTP head are the base model's; config, tokenizer and chat template are identical to Gen1. |
| Languages | English, Simplified Chinese, Traditional Chinese |

## Quickstart

```bash
pip install "sglang==0.5.21" "transformers==5.12.1" requests
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen2 0 8080     # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "m",
  "state": "Customer: my order arrived broken, I want my money back.",
  "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
                "tone": {"type": "choice", "instructions": "Customer tone?",
                         "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
```

- **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen2 --backend hf --port 8080` uses
  transformers only. It is slower and runs bf16.
- **Front end:** `knowline_server.py` is a single file and needs only transformers and requests. It is the same file
  as in Gen1.
- **Full settings:** the exact settings of our Decision Index run are in [INFERENCE.md](INFERENCE.md).

## What changed in Gen2

1. **Finding the gaps.** Gen1 was weakest on knowledge and reasoning (37.7), and behind earlier internal models on
   Jev-style evaluations such as JevBench-hard. The agent traced the latter to the data mix: to balance channels,
   Gen1's data had cut the Jev-style data (OpenJevData and synthetic tasks) by about 70%.
2. **New data (mix F).** On top of Gen1's data:
   - all Jev-style data restored (about 180k OpenJevData rows and 76k synthetic rows);
   - fighting-game data regenerated with balanced behaviour;
   - a new group aimed at Decision Index gaps: ACOS (5% "yes", with hard negatives), WinoGrande XL train, and GSM8K
     train rewritten with Decision Index 0.3 style distractors;
   - the new rows went through decontamination again. Mix F has about 1.05M rows (Gen1's mix E: about 713k).
3. **Training mix F.** Planned for 3 epochs. After the first epoch the model started to memorise the data and its
   Decision Index score declined, so the run was stopped at step 8,455. On their own, mix F checkpoints were below Gen1
   on most evaluations, but clearly more robust to prompt injection.
4. **Weight averaging.** We tried 7 ways of averaging Gen1 with mix F checkpoints:
   - 6 used a single mix F checkpoint (steps 4,500 / 4,800 / 5,000 / 5,500; Gen1 weight 0.4-0.6) and all scored
     61.89-62.14 on the 0.3 public suite;
   - averaging mix F steps 4,500 and 5,000 first, then taking half of that and half of Gen1, scored 62.54 and was at or
     above Gen1 on all of our own evaluations. That is Gen2.
5. **Where the gain comes from.** About 1.35 of the 2.07 points (about two thirds) come from three benchmarks: GSM8K
   (29.5 → 58.0), ACOS (25.0 → 40.8) and WinoGrande (62.1 → 73.8). Mix F added their train splits in Decision Index
   request format. The rest comes from gains across many benchmarks, led by GPQA (16.3 → 23.8), FinEntity, MuSR and
   BPoMP; NLI4CT, iSarcasmEval and When2Call dropped slightly.

## Comparison

Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill,
0.3 new-domain); ours is not available yet.

| model | size | DI 0.3 public | DI 0.3 full | source |
|---|---|---|---|---|
| **KnowLine-4B-Gen2** | 4B | **62.54** | not yet scored | self-run, official kit |
| KnowLine-4B-Gen1 | 4B | 60.47 | not yet scored | self-run, official kit |
| Clef | 27B | 61.71 | 53.08 | board |
| Jev 1.13 | (API) | 57.96 | 60.11 | board |
| RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported |
| ezjev 4B s2 | 4B | 50.82 | 46.95 | board |
| jiwo 4B | 4B | 45.76 | 42.86 | board |
| Nox 4B | 4B | 44.21 | 44.95 | board |

Public-suite scores by area:

| area | KnowLine-4B-Gen2 (0.3 public) | KnowLine-4B-Gen1 (0.3 public) | Jev 1.13 (0.2.1) |
|---|---|---|---|
| Knowledge & reasoning | 42.7 | 37.7 | 51.4 |
| Language | 63.3 | 60.8 | 62.0 |
| Retrieval & routing | 71.3 | 70.9 | 55.4 |
| Tools & agents | 86.6 | 86.8 | 75.1 |
| Arts & taste | 50.2 | 49.5 | 37.7 |

## Releases

The model is trained in a self-evolving loop, and new versions will follow.

| model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes |
|---|---|---|---|---|---|
| [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release |
| [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release |
| KnowLine-4B-Gen2 (this model) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | Gen1 at 0.5 + mix F steps 4,500 and 5,000 at 0.25 each, weight average |
| [KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release (internal run "mix E", step 5,650) |

From Gen2 on we run Decision Index 0.3 only, not 0.2.1.

## Evaluation

All results are self-run and not verified by a third party.

### Held-out and Chinese evaluations

| suite | Gen2 | Gen1 |
|---|---|---|
| In-house evaluation set, English | 69.7 | 68.7 |
| In-house evaluation set, Simplified Chinese | 71.2 | 69.9 |
| In-house evaluation set, Traditional Chinese | 64.8 | 64.1 |
| C-Eval (4 categories, macro) | 78.6 | 77.9 |
| Open-Jev 1.1 test / OOD | 87.4 / 86.6 | 87.4 / 86.6 |
| Prompt injection: answers changed (lower is better) | 4.0% | 8.9% |

### Web operation (Mind2Web official test splits, evaluation only)

| split | Gen2 element selection | Gen2 operation (balanced) | Gen1 element selection | Gen1 operation (balanced) |
|---|---|---|---|---|
| test_task (websites seen in training, new tasks) | 92.9 | 97.0 | 93.1 | 96.6 |
| test_website (new websites) | 90.8 | 97.3 | 90.7 | 96.5 |
| test_domain (new domains) | 91.7 | 97.4 | 91.5 | 98.5 |

- About the same as Gen1; the differences are within noise.
- So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree.
- The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier
  than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard.

### Game harness (KOF '98)

- **Setup:** single-bout character-mirror matches, 18 games per pair, argmax actions; the same settings as Gen1's round
  robin.
- **Result:** 14-3-1 against Jev 1.13, 13-5 against Gen1 and 11-7 against StartLux-Decision-4B; 38-15-1 overall, score
  0.713 [0.58, 0.82].
- **Play style:** much less reliance on the 623C anti-air uppercut, which Gen2 uses 24% of the time (Gen1 about 55% in
  its round robin), with a more varied move mix. It picks moves by distance: uppercut and heavy punch up close,
  special_2 and heavy kick at mid range, and almost only special_2 from far away.
- **Caveats:**
  - 18 games per pair give wide intervals. Gen1's 17-0-1 against Jev in its round robin and Gen2's 14-3-1 here are not
    significantly different.
  - 36 of the 54 games ended at time-out, so most wins are on remaining health rather than by K.O.
  - This is a measured result in this harness, not general fighting-game skill.

### Calibration

Computed on the 0.3 public suite with the method the Decision Index board describes: each field is right or wrong,
confidence is the probability on the chosen option, benchmarks are weighted equally, and there are 10 equal-width bins.

| metric | Gen2 | Gen1 | board median (112 models) | Jev 1.13 |
|---|---|---|---|---|
| ECE (lower is better) | 0.06-0.07 | 0.08-0.09 | 0.084 | 0.074 |
| Brier score (lower is better) | 0.31-0.33 | 0.34-0.36 | 0.49 | 0.36 |
| Confidence ≥95% but wrong | 2.2% | 3.3% | 1.3% | 2.1% |
| Mean confidence / accuracy | 0.83 / 0.76 | 0.84 / 0.74 | | 0.81 / 0.74 |

- **Why ranges:** the board does not say which 32 benchmarks it uses. We checked our computation against three board
  models whose results are public. Our ECE was within about 0.015 of the board's, so we give ranges over the plausible
  benchmark sets.
- **Still overconfident overall:** mean confidence is about 7 points above accuracy. Most of the gap is on hard
  reasoning (HLE, CRUXEval), humour (Humicroedit) and colour judgements (cfcolor).
- **Recommendation:** if you act on probability thresholds, fit a temperature on your own data.

## Disclosures

- **Decision Index format training data:** Gen2 averages models from two training runs. About 25% of mix E and about
  31% of mix F is in Decision Index request format. This includes train splits of public datasets rewritten in that
  format (for example GSM8K, WinoGrande XL and ACOS), and synthetic items written in the same format. No Decision Index
  test item is included.
- **Decontamination:**
  - Every component of a training row of at least 60 characters (a line or paragraph) was checked against the
    components of every Decision Index row and all of our evaluation sets.
  - Rows with a component identical to an evaluation component, or with the same first 50 characters, were removed.
    Text that appears in more than 4 evaluation items counts as boilerplate and is not matched.
  - None of the ~230k rows new in mix F was flagged; the restored Jev-style data comes from an earlier mix that had
    already been decontaminated.
- **Selection on evaluations:** of the 7 averaging variants, we picked the one with the highest Decision Index 0.3
  public score that was also at or above Gen1 on all of our own evaluations.
- **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation
  states.
- **No model outputs as labels:** no output of Jev or any other decision model was used as a training label.
- **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but
  not all were checked by a human.

## Limitations

- **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 42.7 (0.3); MMLU-Pro and HLE are
  still close to the base model.
- **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 58.0; the gain comes from training on GSM8K train
  rewritten in the same format.
- **Prompt injection:** an instruction planted in the state changes the answer about 4% of the time on our injection
  set. Keep untrusted text clearly delimited.
- **Calibration:** overconfident overall; see [Calibration](#calibration).
- **Private tests:** part of our public-suite advantage comes from adapting to the question formats. About two thirds of
  Gen2's gain over Gen1 comes from three benchmarks whose train splits were trained on, so the 0.3 private tests may
  score lower.

## Citation

```bibtex
@misc{knowline4bgen2,
  title  = {KnowLine-4B-Gen2: a 4B decision model},
  author = {PelaAI},
  year   = {2026},
  url    = {https://huggingface.co/PelaAI/KnowLine-4B-Gen2}
}
```

## License

- **Weights:** Apache-2.0, the same as the base model.
- **Code:** `knowline_server.py` is MIT (see the file header).