File size: 7,155 Bytes
64ca6db
 
 
 
602b64b
64ca6db
 
 
14d34c7
64ca6db
14d34c7
64ca6db
 
 
 
 
 
14d34c7
64ca6db
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14d34c7
64ca6db
 
 
 
14d34c7
64ca6db
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
---
license: gemma
base_model: google/gemma-4-E2B-it
language: [en, zh]
tags: [system-one, jev, typed-decisions, calibrated-classification, decision-model, kiosk]
library_name: transformers
pipeline_tag: text-classification
---
# Jevling-E2B-v0.1

**Jevling-E2B-v0.1** is a small *System One* decision model in the family of TypeSafe's Jev: you give it a **state** (any text — a transcript, a ticket, a document) and one or more **typed questions** (choice / yes-no / score), and it answers all of them in **one forward pass, with no text generation**, each as a calibrated probability distribution over the options. It is fine-tuned from `google/gemma-4-E2B-it` for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).

## Quick start (transformers)
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL = "BricksDisplay/jevling-e2b-v0.1"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval()
LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
LETTER_IDS = [tok.encode(c, add_special_tokens=False)[0] for c in LETTERS]

def ask(state, questions):
    """questions: list of dicts {kind: 'choice'|'noul'|'score', text, options, descs (optional)}.
    noul options are always ['no','yes']. Returns one probability list per question (one forward pass)."""
    ids = ([tok.bos_token_id] if tok.bos_token_id is not None else []) + tok.encode(f"<state>\n{state}\n</state>\n", add_special_tokens=False)
    slots, sizes = [], []
    many = len(questions) > 1
    for k, q in enumerate(questions, 1):
        opts = ["no", "yes"] if q["kind"] == "noul" else q["options"]
        tag = {"noul": "yes/no", "score": "score"}.get(q["kind"], "choice")
        text = f"\nQuestion{' '+str(k) if many else ''} ({tag}): {q['text'].strip()}"
        text += "\nLevels:" if q["kind"] == "score" else ("\nOptions:" if q["kind"] == "choice" else "")
        for j, o in enumerate(opts):
            d = (q.get("descs") or [None] * len(opts))[j]
            text += f"\n({LETTERS[j]}) {o}" + (f" — {d}" if d else "")
        text += f"\nAnswer{' '+str(k) if many else ''}: ("
        ids += tok.encode(text, add_special_tokens=False)
        slots.append(len(ids) - 1); sizes.append(len(opts))
    with torch.no_grad():
        logits = model(input_ids=torch.tensor([ids], device=model.device)).logits[0]        # [T, vocab]
    return [torch.softmax(logits[s, LETTER_IDS[:n]].float(), 0).tolist() for s, n in zip(slots, sizes)]

state = "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge."
qs = [
  {"kind": "choice", "text": "Which team should handle this ticket?", "options": ["billing", "technical", "account"],
   "descs": ["payments and refunds", "a product fault", "login or profile settings"]},
  {"kind": "noul", "text": "Is the customer asking for a refund?"},
  {"kind": "score", "text": "How urgent is this?", "options": ["routine", "soon", "urgent", "critical"]},
]
for q, p in zip(qs, ask(state, qs)):
    print(q["text"], [round(x, 3) for x in p])
```
Output for that request:
```
Which team should handle this ticket? [0.998, 0.002, 0.0]     # billing
Is the customer asking for a refund? [0.001, 0.999]            # P(yes) = 1.00
How urgent is this? [0.374, 0.33, 0.2, 0.095]                  # expected level ≈ 1.0 of 0..3
```
Rules of the format: yes/no questions always use the options `no`, `yes`; score questions list ordered levels; give option **descriptions** whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template `system_one`.

## On device
Use the GGUF repo [`BricksDisplay/jevling-e2b-v0.1-GGUF`](https://huggingface.co/BricksDisplay/jevling-e2b-v0.1-GGUF) with the maintained llama.cpp implementation ([`tools/system-one` on mybigday/system-one-llama.cpp, branch `feat/system-one`](https://github.com/mybigday/system-one-llama.cpp/tree/feat/system-one/tools/system-one)). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.

## Evaluation
All numbers are accuracy on datasets the models were **not** trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. Both models of the series are shown; this card's model in bold.

| benchmark | task | Jevling-0.8B-v0.1 | **Jevling-E2B-v0.1** |
|---|---|---|---|
| MASSIVE (en-US) | scenario classification, 18-way | .675 | **.733** |
| BBC News | topic, 5-way | .933 | **.958** |
| TREC | question type, 6-way | .858 | **.850** |
| PAWS | paraphrase yes/no | .508 | **.625** |
| CommitmentBank | NLI, 3-way | .893 | **.875** |
| StrategyQA | yes/no reasoning | .483 | **.567** |
| PubMedQA | yes/no/maybe | .758 | **.667** |
| SciQ | 4-way science QA | .942 | **.975** |
| Social IQa | 3-way | .575 | **.725** |
| TruthfulQA (MC) | multiple choice | .450 | **.633** |
| XStoryCloze (en) | 2-way | .933 | **.958** |
| QuALITY | long-document 4-way QA | .417 | **.500** |
| RewardBench | pairwise preference | .600 | **.817** |
| Hermes function-calling | tool choice | .996 | **.988** |
| Financial PhraseBank | sentiment, 3-way | .608 | **.658** |
| JevBench easy / original / hard (231 items) | typed decisions | 1.000 / .833 / .441 | **1.000 / .903 / .441** |
| zh-TW kiosk set (ours, **synthetic-derived**, 255 states) | intent acc / completeness AUROC / is-order / noise / size | .969 / .971 / .996 / 1.000 / 1.000 | **.973 / .989 / .995 / 1.000 / 1.000** |

JevBench *hard* (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 — the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.

## Limitations
- Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.5); compute arithmetic in code and put the result in the state.
- Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
- When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
- Not a chat model: it does not generate text.

## Training data
Fine-tuned on a mix of public classification / QA / preference / tool-use datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.


## Licence and release status
**v0.1 is a research / non-commercial release.** Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).