Text Classification
Transformers
Safetensors
English
Chinese
qwen3_5
image-text-to-text
decision-model
system-one
decision-index
lora-merged
Instructions to use PelaAI/KnowLine-4B-Gen3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen3")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen3") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from PelaAI/KnowLine-4B-Gen3: direct link, hf CLI and curl.
- Browser
- Download file 13.2 kB
-
https://huggingface.co/PelaAI/KnowLine-4B-Gen3/resolve/main/README.md
- Command line
-
hf download hf://PelaAI/KnowLine-4B-Gen3/README.md
-
curl -L -o README.md https://huggingface.co/PelaAI/KnowLine-4B-Gen3/resolve/main/README.md
13.2 kB
| license: apache-2.0 | |
| base_model: | |
| - Qwen/Qwen3.5-4B | |
| base_model_relation: finetune | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| language: | |
| - en | |
| - zh | |
| tags: | |
| - decision-model | |
| - system-one | |
| - decision-index | |
| - lora-merged | |
| # KnowLine-4B-Gen3 | |
| **The third release of PelaAI's 4B System One decision model, trained on a single all-in-one machine by an agent-driven | |
| data loop.** | |
| [Inference guide](INFERENCE.md) · [中文说明](README.zh.md) · weights Apache-2.0 · previous version: | |
| [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) | |
| Highlights: | |
| - **Decision Index 0.3, public suite: 63.11** (self-run), 0.57 above Gen2 (62.54). For comparison: Perplexity Decider | |
| v1.1 (27B) scores 62.25 and Jev 1.13 57.96 (board), and the best ≤5B model on the board has a public score of 50.82; | |
| see [Comparison](#comparison). | |
| - **Better than Gen2 on our Jev-style and Chinese evaluations:** Open-Jev 1.1 test / OOD 88.4 / 87.2 (Gen2: 87.4 / | |
| 86.6), C-Eval 79.2 (78.6). Instructions planted in the state change the answer 3.4% of the time, down from 4.0%. | |
| - **Calibration:** computed the board's way, the Brier score is about 0.31-0.33, roughly 3rd of 114 models on the | |
| board, and ECE is about 0.07-0.08 (board median 0.084); see [Calibration](#calibration). | |
| - **KOF '98 harness: 16-1-1 against Jev 1.13 and 12-4-2 against StartLux-Decision-4B** in single-bout mirror matches; see | |
| [Game harness](#game-harness-kof-98). | |
| - **An AI agent ran the whole loop.** In each round it: | |
| - found the areas where the model was weak; | |
| - proposed a targeted data group; | |
| - built and decontaminated the data; | |
| - trained and evaluated on it; | |
| - kept or rejected the change based on the evidence. | |
| For this round, see [What changed in Gen3](#what-changed-in-gen3). | |
| KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev. | |
| ## What it does | |
| You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model | |
| answers every question with a probability distribution over its options: | |
| - one forward pass per question, with no generated text; | |
| - the output is the probability of each option's label token; | |
| - existing Jev clients only need a new base URL; Gen1 and Gen2 users only change the model name, since the interface | |
| and serving settings are the same. | |
| | | | | |
| |---|---| | |
| | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0) | | |
| | Training | LoRA SFT (rank 32, alpha 64, language model only), merged into bf16 weights. | | |
| | Release format | bf16 weights. The vision tower and MTP head are the base model's; config, tokenizer and chat template are identical to Gen1 and Gen2. | | |
| | Languages | English, Simplified Chinese, Traditional Chinese | | |
| ## Quickstart | |
| ```bash | |
| pip install "sglang==0.5.21" "transformers==5.12.1" requests | |
| bash serve_knowline.sh PelaAI/KnowLine-4B-Gen3 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080 | |
| curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{ | |
| "model": "m", | |
| "state": "Customer: my order arrived broken, I want my money back.", | |
| "questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"}, | |
| "tone": {"type": "choice", "instructions": "Customer tone?", | |
| "criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}' | |
| ``` | |
| - **Without SGLang:** `python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend hf --port 8080` uses | |
| transformers only. It is slower and runs bf16. | |
| - **Front end:** `knowline_server.py` is a single file and needs only transformers and requests. It is the same file | |
| as in Gen2. | |
| - **Full settings:** the exact settings of our Decision Index run are in [INFERENCE.md](INFERENCE.md). | |
| ## What changed in Gen3 | |
| 1. **Finding the gaps.** Gen2 was still weakest on knowledge and reasoning (42.7). | |
| 2. **New data.** On top of the previous round's data, a knowledge replay group: multiple-choice items from public train | |
| splits covering general knowledge, science, medicine, math word problems and logic, in English and Chinese. None of | |
| them comes from a Decision Index benchmark, and the new rows went through decontamination again. | |
| 3. **Training.** One epoch, with the learning rate decayed all the way down. The replay raised knowledge and reasoning a | |
| little but cost BPoMP in the arts area. | |
| 4. **Selection.** This round compared 13 candidates, which scored 62.54-63.25 on the 0.3 public suite. Gen3 scores | |
| 63.11 on that suite, within noise of the highest score. Of the 13 candidates it was the best on our Jev-style | |
| evaluations (JevBench 86.6, JevBench-hard 72.1) and C-Eval (79.2), and near the top on out-of-distribution | |
| generalisation (Open-Jev 1.1 OOD 87.2). | |
| 5. **Where the gain comes from.** The gain over Gen2 is spread over many benchmarks, led by NLI4CT (57.6 → 62.3), | |
| CRUXEval (38.8 → 44.0), WinoGrande (73.8 → 78.8), RAGTruth (53.3 → 56.8) and GSM8K (58.0 → 61.5). The train splits of | |
| WinoGrande and GSM8K are in the training data in Decision Index request format. BPoMP (59.4 → 54.5) and PhishNChips | |
| (95.0 → 92.0) dropped. | |
| ## Comparison | |
| Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill, | |
| 0.3 new-domain); ours is not available yet. | |
| | model | size | DI 0.3 public | DI 0.3 full | source | | |
| |---|---|---|---|---| | |
| | **KnowLine-4B-Gen3** | 4B | **63.11** | not yet scored | self-run, official kit | | |
| | KnowLine-4B-Gen2 | 4B | 62.54 | not yet scored | self-run, official kit | | |
| | KnowLine-4B-Gen1 | 4B | 60.47 | not yet scored | self-run, official kit | | |
| | Perplexity Decider v1.1 | 27B | 62.25 | 62.75 | board | | |
| | Clef | 27B | 61.71 | 53.08 | board | | |
| | Jev 1.13 | (API) | 57.96 | 60.11 | board | | |
| | RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported | | |
| | ezjev 4B s2 | 4B | 50.82 | 46.95 | board | | |
| | jiwo 4B | 4B | 45.76 | 42.86 | board | | |
| | Nox 4B | 4B | 44.21 | 44.95 | board | | |
| Public-suite scores by area (all 0.3 public): | |
| | area | KnowLine-4B-Gen3 | KnowLine-4B-Gen2 | Perplexity Decider v1.1 (27B) | Jev 1.13 | | |
| |---|---|---|---|---| | |
| | Knowledge & reasoning | 43.9 | 42.7 | 52.0 | 53.9 | | |
| | Language | 64.8 | 63.3 | 67.5 | 59.2 | | |
| | Retrieval & routing | 70.9 | 71.3 | 61.3 | 55.4 | | |
| | Tools & agents | 86.6 | 86.6 | 78.9 | 75.1 | | |
| | Arts & taste | 49.8 | 50.2 | 47.0 | 39.1 | | |
| ## Releases | |
| The model is trained in a self-evolving loop, and new versions will follow. | |
| | model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes | | |
| |---|---|---|---|---|---| | |
| | [KnowLine-4B-Gen4](https://huggingface.co/PelaAI/KnowLine-4B-Gen4) | 2026-10-10 | 64.90 | — | 69.7 / 70.1 / 65.1 | fourth release | | |
| | KnowLine-4B-Gen3 (this model) | 2026-10-09 | 63.11 | — | 69.4 / 70.8 / 65.0 | third release | | |
| | [KnowLine-4B-Gen2](https://huggingface.co/PelaAI/KnowLine-4B-Gen2) | 2026-10-08 | 62.54 | — | 69.7 / 71.2 / 64.8 | second release | | |
| | [KnowLine-4B-Gen1](https://huggingface.co/PelaAI/KnowLine-4B-Gen1) | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release | | |
| From Gen2 on we run Decision Index 0.3 only, not 0.2.1. | |
| ## Evaluation | |
| All results are self-run and not verified by a third party. | |
| ### Held-out and Chinese evaluations | |
| | suite | Gen3 | Gen2 | | |
| |---|---|---| | |
| | In-house evaluation set, English | 69.4 | 69.7 | | |
| | In-house evaluation set, Simplified Chinese | 70.8 | 71.2 | | |
| | In-house evaluation set, Traditional Chinese | 65.0 | 64.8 | | |
| | C-Eval (4 categories, macro) | 79.2 | 78.6 | | |
| | Open-Jev 1.1 test / OOD | 88.4 / 87.2 | 87.4 / 86.6 | | |
| | Prompt injection: answers changed (lower is better) | 3.4% | 4.0% | | |
| ### Web operation (Mind2Web official test splits, evaluation only) | |
| | split | Gen3 element selection | Gen3 operation (balanced) | Gen2 element selection | Gen2 operation (balanced) | | |
| |---|---|---|---|---| | |
| | test_task (websites seen in training, new tasks) | 92.7 | 97.0 | 92.9 | 97.0 | | |
| | test_website (new websites) | 91.3 | 97.3 | 90.8 | 97.3 | | |
| | test_domain (new domains) | 91.8 | 97.2 | 91.7 | 97.4 | | |
| - About the same as Gen2; the differences are within noise. | |
| - So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree. | |
| - The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier | |
| than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard. | |
| ### Game harness (KOF '98) | |
| - **Setup:** single-bout character-mirror matches, 18 games per pair, argmax actions; the same settings as Gen1's round | |
| robin and Gen2's matches. | |
| - **Result:** 16-1-1 against Jev 1.13, 12-4-2 against StartLux-Decision-4B and 4-14 against Gen2; 32-19-3 overall, | |
| score 0.620 [0.49, 0.74]. | |
| - **Play style:** it picks moves by distance: mostly the 623C anti-air uppercut up close (75%), special_2 and heavy kick | |
| at mid range, and almost only special_2 from far away. Overall it uses the uppercut 23% of the time. | |
| - **Caveats:** | |
| - 18 games per pair give wide intervals. | |
| - 37 of the 54 games ended at time-out, so most wins are on remaining health. Gen3 won 8 games by K.O. (2 against | |
| Jev, 4 against StartLux, 2 against Gen2). | |
| - This is a measured result in this harness, not general fighting-game skill. | |
| ### Calibration | |
| Computed on the 0.3 public suite with the method the Decision Index board describes: each field is right or wrong, | |
| confidence is the probability on the chosen option, benchmarks are weighted equally, and there are 10 equal-width bins. | |
| | metric | Gen3 | Gen2 | board median (114 models) | Jev 1.13 | | |
| |---|---|---|---|---| | |
| | ECE (lower is better) | 0.07-0.08 | 0.06-0.07 | 0.084 | 0.074 | | |
| | Brier score (lower is better) | 0.31-0.33 | 0.31-0.33 | 0.49 | 0.36 | | |
| | Confidence ≥95% but wrong | 2.7% | 2.2% | 1.3% | 2.1% | | |
| | Mean confidence / accuracy | 0.84 / 0.76 | 0.83 / 0.76 | | 0.81 / 0.74 | | |
| - **Why ranges:** the board does not say which 32 benchmarks it uses. We checked our computation against three board | |
| models whose results are public. Our ECE was within about 0.015 of the board's, so we give ranges over the plausible | |
| benchmark sets. | |
| - **Still overconfident overall:** mean confidence is about 7-8 points above accuracy. Most of the gap is on hard | |
| reasoning (HLE, CRUXEval, MuSR) and humour (Humicroedit, New Yorker caption matching). | |
| - **Recommendation:** if you act on probability thresholds, fit a temperature on your own data. | |
| ## Disclosures | |
| - **Decision Index format training data:** Gen3 comes from three training rounds. In each, about 25-31% of the data is | |
| in Decision Index request format. This includes train splits of public datasets | |
| rewritten in that format (for example GSM8K, WinoGrande XL and ACOS), and synthetic items written in the same format. | |
| No Decision Index test item is included. | |
| - **Decontamination:** | |
| - Every component of a training row of at least 60 characters (a line or paragraph) was checked against the | |
| components of every Decision Index row and all of our evaluation sets. | |
| - Rows with a component identical to an evaluation component, or with the same first 50 characters, were removed. | |
| Text that appears in more than 4 evaluation items counts as boilerplate and is not matched. | |
| - None of the rows new in this round was flagged; the rest of the data had already been decontaminated in earlier | |
| rounds. | |
| - **Selection on evaluations:** of the 13 averaging variants, we picked one within noise of the highest Decision Index | |
| 0.3 public score that was also best on our Jev-style evaluations. | |
| - **Game data:** labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation | |
| states. | |
| - **No model outputs as labels:** no output of Jev or any other decision model was used as a training label. | |
| - **Teacher-labelled synthetic data:** LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but | |
| not all were checked by a human. | |
| ## Limitations | |
| - **Knowledge-heavy reasoning:** weaker than larger models. The knowledge area is 43.9 (0.3); MMLU-Pro moved little | |
| (51.7) and HLE stays at chance level. | |
| - **Math:** answered without reasoning. The rebuilt 0.3 GSM8K scores 61.5; the gain comes from training on GSM8K train | |
| rewritten in the same format. | |
| - **Prompt injection:** an instruction planted in the state changes the answer about 3% of the time on our injection | |
| set. Keep untrusted text clearly delimited. | |
| - **Calibration:** overconfident overall; see [Calibration](#calibration). | |
| - **Private tests:** part of our public-suite advantage comes from adapting to the question formats. Some of Gen3's | |
| largest gains over Gen2 are on benchmarks whose train splits were trained on, so the 0.3 private tests may score | |
| lower. | |
| ## Citation | |
| ```bibtex | |
| @misc{knowline4bgen3, | |
| title = {KnowLine-4B-Gen3: a 4B decision model}, | |
| author = {PelaAI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen3} | |
| } | |
| ``` | |
| ## License | |
| - **Weights:** Apache-2.0, the same as the base model. | |
| - **Code:** `knowline_server.py` is MIT (see the file header). | |