Text Classification
Transformers
Safetensors
English
Chinese
qwen3_5
image-text-to-text
decision-model
system-one
decision-index
lora-merged
model-soup
Instructions to use PelaAI/KnowLine-4B-Gen2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen2")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen2") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from PelaAI/KnowLine-4B-Gen2: direct link, hf CLI and curl.
- Browser
- Download file 14.1 kB
-
https://huggingface.co/PelaAI/KnowLine-4B-Gen2/resolve/main/README.md
- Command line
-
hf download hf://PelaAI/KnowLine-4B-Gen2/README.md
-
curl -L -o README.md https://huggingface.co/PelaAI/KnowLine-4B-Gen2/resolve/main/README.md
14.1 kB
| license: apache-2.0 | |
| base_model: | |
| - Qwen/Qwen3.5-4B | |
| base_model_relation: finetune | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| language: | |
| - en | |
| - zh | |
| tags: | |
| - decision-model | |
| - system-one | |
| - decision-index | |
| - lora-merged | |
| - model-soup | |
| # KnowLine-4B-Gen2 | |
| **The second release of PelaAI's 4B System One decision model: a weight average of KnowLine-4B-Gen1 and the model from | |
| the next training round.** | |
| [Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0 · previous version: | |
| [KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1) | |
| > **Newer version:** [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) is out, with a Decision Index 0.3 public score of | |
| > 63.11. | |
| Highlights: | |
| - **Decision Index 0.3, public suite: 62.54** (self-run), 2.07 above Gen1 (60.47). For comparison: Jev 1.13 scores | |
| 57.96 (board), and the best ≤5B model on the board has a public score of 50.82; see [Comparison](#comparison). | |
| - **At or above Gen1 on all of our own evaluations:** in-house held-out set (en / zh-Hans / zh-Hant) 69.7 / 71.2 / 64.8 | |
| (Gen1: 68.7 / 69.9 / 64.1), C-Eval 78.6 (77.9). Instructions planted in the state change the answer 4.0% of the | |
| time, down from 8.9%. | |
| - **Better calibrated:** computed the board's way, ECE is about 0.06-0.07 (Gen1 about 0.08-0.09, board median 0.084), | |
| and the Brier score is about 0.31-0.33, roughly 3rd of 113 models on the board; see [Calibration](#calibration). | |
| - **KOF '98 harness: 14-3-1 against Jev 1.13 and 13-5 against Gen1** in single-bout mirror matches; see | |
| [Game harness](#game-harness-kof-98). | |
| - **An AI agent ran the whole loop.** In each round it: | |
| - found the areas where the model was weak; | |
| - proposed a targeted data group; | |
| - built and decontaminated the data; | |
| - trained and evaluated on it; | |
| - kept or rejected the change based on the evidence. | |
| For this round, see [What changed in Gen2](#what-changed-in-gen2). | |
| KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev. | |
| ## What it does | |
| You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model | |
| answers every question with a probability distribution over its options: | |
| - one forward pass per question, with no generated text; | |
| - the output is the probability of each option's label token; | |
| - existing Jev clients only need a new base URL; Gen1 users only change the model name, since the interface and | |
| serving settings are the same. | |
| | | | | |
| |---|---| | |
| | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) | | |
| | Training | Two LoRA SFT runs (rank 32, alpha 64, language model only), merged, then averaged tensor by tensor: Gen1 (run "mix E", step 5,650) at 0.5, and steps 4,500 and 5,000 of run "mix F" at 0.25 each. | | |
| | Release format | bf16 weights. The vision tower and MTP head are the base model's; config, tokenizer and chat template are identical to Gen1. | | |
| | Languages | English, Simplified Chinese, Traditional Chinese | | |
| ## Quickstart | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests | |
| bash serve_knowline.sh PelaAI/KnowLine-4B-Gen2 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080 | |
| curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{ | |
| "model": "m", | |
| "state": "Customer: my order arrived broken, I want my money back.", | |
| "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"}, | |
| "tone": {"type": "choice", "instructions": "Customer tone?", | |
| "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}' | |
| ``` | |
| - **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen2 --backend hf --port 8080` uses | |
| transformers only. It is slower and runs bf16. | |
| - **Front end:** `knowline_server.py` is a single file and needs only transformers and requests. It is the same file | |
| as in Gen1. | |
| - **Full settings:** the exact settings of our Decision Index run are in [INFERENCE.md](INFERENCE.md). | |
| ## What changed in Gen2 | |
| 1. **Finding the gaps.** Gen1 was weakest on knowledge and reasoning (37.7), and behind earlier internal models on | |
| Jev-style evaluations such as JevBench-hard. The agent traced the latter to the data mix: to balance channels, | |
| Gen1's data had cut the Jev-style data (OpenJevData and synthetic tasks) by about 70%. | |
| 2. **New data (mix F).** On top of Gen1's data: | |
| - all Jev-style data restored (about 180k OpenJevData rows and 76k synthetic rows); | |
| - fighting-game data regenerated with balanced behaviour; | |
| - a new group aimed at Decision Index gaps: ACOS (5% "yes", with hard negatives), WinoGrande XL train, and GSM8K | |
| train rewritten with Decision Index 0.3 style distractors; | |
| - the new rows went through decontamination again. Mix F has about 1.05M rows (Gen1's mix E: about 713k). | |
| 3. **Training mix F.** Planned for 3 epochs. After the first epoch the model started to memorise the data and its | |
| Decision Index score declined, so the run was stopped at step 8,455. On their own, mix F checkpoints were below Gen1 | |
| on most evaluations, but clearly more robust to prompt injection. | |
| 4. **Weight averaging.** We tried 7 ways of averaging Gen1 with mix F checkpoints: | |
| - 6 used a single mix F checkpoint (steps 4,500 / 4,800 / 5,000 / 5,500; Gen1 weight 0.4-0.6) and all scored | |
| 61.89-62.14 on the 0.3 public suite; | |
| - averaging mix F steps 4,500 and 5,000 first, then taking half of that and half of Gen1, scored 62.54 and was at or | |
| above Gen1 on all of our own evaluations. That is Gen2. | |
| 5. **Where the gain comes from.** About 1.35 of the 2.07 points (about two thirds) come from three benchmarks: GSM8K | |
| (29.5 → 58.0), ACOS (25.0 → 40.8) and WinoGrande (62.1 → 73.8). Mix F added their train splits in Decision Index | |
| request format. The rest comes from gains across many benchmarks, led by GPQA (16.3 → 23.8), FinEntity, MuSR and | |
| BPoMP; NLI4CT, iSarcasmEval and When2Call dropped slightly. | |
| ## Comparison | |
| Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill, | |
| 0.3 new-domain); ours is not available yet. | |
| | model | size | DI 0.3 public | DI 0.3 full | source | | |
| |---|---|---|---|---| | |
| | **KnowLine-4B-Gen2** | 4B | **62.54** | not yet scored | self-run, official kit | | |
| | KnowLine-4B-Gen1 | 4B | 60.47 | not yet scored | self-run, official kit | | |
| | Clef | 27B | 61.71 | 53.08 | board | | |
| | Jev 1.13 | (API) | 57.96 | 60.11 | board | | |
| | RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported | | |
| | ezjev 4B s2 | 4B | 50.82 | 46.95 | board | | |
| | jiwo 4B | 4B | 45.76 | 42.86 | board | | |
| | Nox 4B | 4B | 44.21 | 44.95 | board | | |
| Public-suite scores by area: | |
| | area | KnowLine-4B-Gen2 (0.3 public) | KnowLine-4B-Gen1 (0.3 public) | Jev 1.13 (0.2.1) | | |
| |---|---|---|---| | |
| | Knowledge & reasoning | 42.7 | 37.7 | 51.4 | | |
| | Language | 63.3 | 60.8 | 62.0 | | |
| | Retrieval & routing | 71.3 | 70.9 | 55.4 | | |
| | Tools & agents | 86.6 | 86.8 | 75.1 | | |
| | Arts & taste | 50.2 | 49.5 | 37.7 | | |
| ## Releases | |
| The model is trained in a self-evolving loop, and new versions will follow. | |
| | model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes | | |
| |---|---|---|---|---|---| | |
| | [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release | | |
| | [KnowLine-4B-Gen3](https://huggingface.co/PelaAI/KnowLine-4B-Gen3) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release | | |
| | KnowLine-4B-Gen2 (this model) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | Gen1 at 0.5 + mix F steps 4,500 and 5,000 at 0.25 each, weight average | | |
| | [KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release (internal run "mix E", step 5,650) | | |
| From Gen2 on we run Decision Index 0.3 only, not 0.2.1. | |
| ## Evaluation | |
| All results are self-run and not verified by a third party. | |
| ### Held-out and Chinese evaluations | |
| | suite | Gen2 | Gen1 | | |
| |---|---|---| | |
| | In-house evaluation set, English | 69.7 | 68.7 | | |
| | In-house evaluation set, Simplified Chinese | 71.2 | 69.9 | | |
| | In-house evaluation set, Traditional Chinese | 64.8 | 64.1 | | |
| | C-Eval (4 categories, macro) | 78.6 | 77.9 | | |
| | Open-Jev 1.1 test / OOD | 87.4 / 86.6 | 87.4 / 86.6 | | |
| | Prompt injection: answers changed (lower is better) | 4.0% | 8.9% | | |
| ### Web operation (Mind2Web official test splits, evaluation only) | |
| | split | Gen2 element selection | Gen2 operation (balanced) | Gen1 element selection | Gen1 operation (balanced) | | |
| |---|---|---|---|---| | |
| | test_task (websites seen in training, new tasks) | 92.9 | 97.0 | 93.1 | 96.6 | | |
| | test_website (new websites) | 90.8 | 97.3 | 90.7 | 96.5 | | |
| | test_domain (new domains) | 91.7 | 97.4 | 91.5 | 98.5 | | |
| - About the same as Gen1; the differences are within noise. | |
| - So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree. | |
| - The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier | |
| than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard. | |
| ### Game harness (KOF '98) | |
| - **Setup:** single-bout character-mirror matches, 18 games per pair, argmax actions; the same settings as Gen1's round | |
| robin. | |
| - **Result:** 14-3-1 against Jev 1.13, 13-5 against Gen1 and 11-7 against StartLux-Decision-4B; 38-15-1 overall, score | |
| 0.713 [0.58, 0.82]. | |
| - **Play style:** much less reliance on the 623C anti-air uppercut, which Gen2 uses 24% of the time (Gen1 about 55% in | |
| its round robin), with a more varied move mix. It picks moves by distance: uppercut and heavy punch up close, | |
| special_2 and heavy kick at mid range, and almost only special_2 from far away. | |
| - **Caveats:** | |
| - 18 games per pair give wide intervals. Gen1's 17-0-1 against Jev in its round robin and Gen2's 14-3-1 here are not | |
| significantly different. | |
| - 36 of the 54 games ended at time-out, so most wins are on remaining health rather than by K.O. | |
| - This is a measured result in this harness, not general fighting-game skill. | |
| ### Calibration | |
| Computed on the 0.3 public suite with the method the Decision Index board describes: each field is right or wrong, | |
| confidence is the probability on the chosen option, benchmarks are weighted equally, and there are 10 equal-width bins. | |
| | metric | Gen2 | Gen1 | board median (112 models) | Jev 1.13 | | |
| |---|---|---|---|---| | |
| | ECE (lower is better) | 0.06-0.07 | 0.08-0.09 | 0.084 | 0.074 | | |
| | Brier score (lower is better) | 0.31-0.33 | 0.34-0.36 | 0.49 | 0.36 | | |
| | Confidence ≥95% but wrong | 2.2% | 3.3% | 1.3% | 2.1% | | |
| | Mean confidence / accuracy | 0.83 / 0.76 | 0.84 / 0.74 | | 0.81 / 0.74 | | |
| - **Why ranges:** the board does not say which 32 benchmarks it uses. We checked our computation against three board | |
| models whose results are public. Our ECE was within about 0.015 of the board's, so we give ranges over the plausible | |
| benchmark sets. | |
| - **Still overconfident overall:** mean confidence is about 7 points above accuracy. Most of the gap is on hard | |
| reasoning (HLE, CRUXEval), humour (Humicroedit) and colour judgements (cfcolor). | |
| - **Recommendation:** if you act on probability thresholds, fit a temperature on your own data. | |
| ## Disclosures | |
| - **Decision Index format training data:** Gen2 averages models from two training runs. About 25% of mix E and about | |
| 31% of mix F is in Decision Index request format. This includes train splits of public datasets rewritten in that | |
| format (for example GSM8K, WinoGrande XL and ACOS), and synthetic items written in the same format. No Decision Index | |
| test item is included. | |
| - **Decontamination:** | |
| - Every component of a training row of at least 60 characters (a line or paragraph) was checked against the | |
| components of every Decision Index row and all of our evaluation sets. | |
| - Rows with a component identical to an evaluation component, or with the same first 50 characters, were removed. | |
| Text that appears in more than 4 evaluation items counts as boilerplate and is not matched. | |
| - None of the ~230k rows new in mix F was flagged; the restored Jev-style data comes from an earlier mix that had | |
| already been decontaminated. | |
| - **Selection on evaluations:** of the 7 averaging variants, we picked the one with the highest Decision Index 0.3 | |
| public score that was also at or above Gen1 on all of our own evaluations. | |
| - **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation | |
| states. | |
| - **No model outputs as labels:** no output of Jev or any other decision model was used as a training label. | |
| - **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but | |
| not all were checked by a human. | |
| ## Limitations | |
| - **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 42.7 (0.3); MMLU-Pro and HLE are | |
| still close to the base model. | |
| - **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 58.0; the gain comes from training on GSM8K train | |
| rewritten in the same format. | |
| - **Prompt injection:** an instruction planted in the state changes the answer about 4% of the time on our injection | |
| set. Keep untrusted text clearly delimited. | |
| - **Calibration:** overconfident overall; see [Calibration](#calibration). | |
| - **Private tests:** part of our public-suite advantage comes from adapting to the question formats. About two thirds of | |
| Gen2's gain over Gen1 comes from three benchmarks whose train splits were trained on, so the 0.3 private tests may | |
| score lower. | |
| ## Citation | |
| ```bibtex | |
| @misc{knowline4bgen2, | |
| title = {KnowLine-4B-Gen2: a 4B decision model}, | |
| author = {PelaAI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen2} | |
| } | |
| ``` | |
| ## License | |
| - **Weights:** Apache-2.0, the same as the base model. | |
| - **Code:** `knowline_server.py` is MIT (see the file header). | |