KnowLine-4B-Gen1 / README.md
PEScn's picture
README.md: link KnowLine-4B-Gen4
47a81a0 verified
|
Raw History Blame Contribute Delete
8.76 kB
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-4B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-classification
language:
- en
- zh
tags:
- decision-model
- system-one
- decision-index
- lora-merged
---
# KnowLine-4B-Gen1
**A 4B System One decision model from PelaAI, trained on a single all-in-one machine by an agent-driven data loop.**
[Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0
> **Newer versions:** [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) (Decision Index 0.3 public score 63.11) and
> [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) (62.54) are out. Both are at or above this model on our held-out, C-Eval, Open-Jev and prompt-injection
> evaluations.
Highlights:
- **Decision Index 0.3, public suite: 60.47** (self-run). For comparison: Jev 1.13 scores 57.96 (board), and the best
≤5B model on the board has a public score of 50.82; see [Comparison](#comparison).
- **An AI agent ran the whole loop.** In each round it:
- found the areas where the model was weak;
- proposed a targeted data group;
- built and decontaminated the data;
- trained and evaluated on it;
- kept or rejected the change based on the evidence.
- **KOF '98 harness: 17 wins, 0 losses, 1 draw against Jev 1.13** in a single-bout mirror round robin. For the play
style, see [Game harness](#game-harness-kof-98).
KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev.
## What it does
You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model
answers every question with a probability distribution over its options:
- one forward pass per question, with no generated text;
- the output is the probability of each option's label token;
- existing Jev clients only need a new base URL.
| | |
|---|---|
| Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) |
| Training | LoRA SFT (rank 32, alpha 64) on the language model, one epoch over about 713k rows. This release is step 5,650 of 5,655, the lowest validation loss. |
| Release format | LoRA merged into the base; bf16 weights. The vision tower and MTP head are unchanged. |
| Languages | English, Simplified Chinese, Traditional Chinese |
## Quickstart
```bash
pip install "sglang==0.5.21" "transformers==5.12.1" requests
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen1 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "m",
"state": "Customer: my order arrived broken, I want my money back.",
"questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
"tone": {"type": "choice", "instructions": "Customer tone?",
"criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
```
- **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen1 --backend hf --port 8080` uses
transformers only. It is slower and runs bf16.
- **Front end:** `knowline_server.py` is a single file and needs only transformers and requests.
- **Full settings:** the exact settings of our Decision Index runs are in [INFERENCE.md](INFERENCE.md).
## Comparison
Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill,
0.3 new-domain); ours is not available yet.
| model | size | DI 0.3 public | DI 0.3 full | source |
|---|---|---|---|---|
| **KnowLine-4B-Gen1** | 4B | **60.47** | not yet scored | self-run, official kit |
| Clef | 27B | 61.71 | 53.08 | board |
| Jev 1.13 | (API) | 57.96 | 60.11 | board |
| RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported |
| ezjev 4B s2 | 4B | 50.82 | 46.95 | board |
| jiwo 4B | 4B | 45.76 | 42.86 | board |
| Nox 4B | 4B | 44.21 | 44.95 | board |
Our public-suite scores by area:
| area | KnowLine-4B-Gen1 (0.3 public) | Jev 1.13 (0.2.1) |
|---|---|---|
| Knowledge & reasoning | 37.7 | 51.4 |
| Language | 60.8 | 62.0 |
| Retrieval & routing | 70.9 | 55.4 |
| Tools & agents | 86.8 | 75.1 |
| Arts & taste | 49.5 | 37.7 |
## Releases
The model is trained in a self-evolving loop, and new versions will follow.
| model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes |
|---|---|---|---|---|---|
| [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release |
| [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release |
| [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | this model at 0.5 + mix F steps 4,500 and 5,000 at 0.25 each, weight average |
| KnowLine-4B-Gen1 (this model) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release (internal run "mix E", step 5,650) |
## Evaluation
All results are self-run and not verified by a third party.
### Held-out and Chinese evaluations
| suite | accuracy |
|---|---|
| In-house evaluation set, English | 68.7 |
| In-house evaluation set, Simplified Chinese | 69.9 |
| In-house evaluation set, Traditional Chinese | 64.1 |
| C-Eval (4 categories, macro) | 77.9 |
| Open-Jev 1.1 test / OOD | 87.4 / 86.6 |
### Web operation (Mind2Web official test splits, evaluation only)
| split | element selection | operation (balanced) |
|---|---|---|
| test_task (websites seen in training, new tasks) | 93.1 | 96.6 |
| test_website (new websites) | 90.7 | 96.5 |
| test_domain (new domains) | 91.5 | 98.5 |
- So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree.
- The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier
than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard.
### Game harness (KOF '98)
- **Setup:** single-bout character-mirror round robin, 11 players, 18 games per pair (990 games), argmax actions.
- **Result:** win rate 0.883 [0.83, 0.92], Elo 1910, tied for first of 11 with another internal checkpoint.
- **Head-to-head:** 17-0-1 against Jev 1.13, 15-3 against StartLux-Decision-4B, 18-0 against Clef-Flash.
- **Caveat:** the policy relies heavily on one move, a 623C anti-air uppercut used about 55% of the time. This is a
measured result in this harness, not general fighting-game skill.
## Disclosures
- **Decision Index training data:** about 25% of the data is in Decision Index format, that is, synthetic items written
in the benchmarks' request formats. No Decision Index test item is included.
- **Decontamination:**
- Every component text (state, instructions, option texts) was checked against every Decision Index row and all of
our evaluation sets.
- Rows that are identical, or share any run of 50 consecutive characters, were removed.
- **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation
states.
- **No model outputs as labels:** no output of Jev or any other decision model was used as a training label.
- **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but
not all were checked by a human.
## Limitations
- **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 37.7 (0.3); MMLU-Pro, GPQA and HLE are
close to the base model.
- **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 29.5.
- **Prompt injection:** an instruction planted in the state changes the answer about 9% of the time on our injection
set. Keep untrusted text clearly delimited.
- **Calibration:** ECE is about 0.05-0.06 on choice and yes/no questions and about 0.12 on score questions. For score
questions, fit a temperature on your own data. On the Decision Index 0.3 public suite, computed the board's way, ECE
is about 0.08-0.09 and 3.3% of answers are wrong at ≥95% confidence (board median 1.3%): the model is overconfident
overall.
- **Private tests:** part of our public-suite advantage comes from adapting to the question formats, so the 0.3 private
tests may score lower.
## Citation
```bibtex
@misc{knowline4bgen1,
title = {KnowLine-4B-Gen1: a 4B decision model},
author = {PelaAI},
year = {2026},
url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen1}
}
```
## License
- **Weights:** Apache-2.0, the same as the base model.
- **Code:** `knowline_server.py` is MIT (see the file header).