--- license: apache-2.0 base_model: google/gemma-4-E2B-it language: [en, zh] tags: [system-one, jev, typed-decisions, calibrated-classification, kiosk] library_name: transformers pipeline_tag: text-classification --- # Jevling-E2B-v1 **Jevling-E2B-v1** is a small *System One* decision model in the family of TypeSafe's Jev: you give it a **state** (any text — a transcript, a ticket, a document) and one or more **typed questions** (choice / yes-no / score), and it answers all of them in **one forward pass, with no text generation**, each as a calibrated probability distribution over the options. It is fine-tuned from `google/gemma-4-E2B-it` for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering). ## Quick start (transformers) ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM MODEL = "BricksDisplay/jevling-e2b-v1" tok = AutoTokenizer.from_pretrained(MODEL) model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval() # on ROCm add attn_implementation="eager" LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" LETTER_IDS = [tok.encode(c, add_special_tokens=False)[0] for c in LETTERS] def ask(state, questions): """questions: list of dicts {kind: 'choice'|'noul'|'score', text, options, descs (optional)}. noul options are always ['no','yes']. Returns one probability list per question (one forward pass).""" ids = ([tok.bos_token_id] if tok.bos_token_id is not None else []) + tok.encode(f"\n{state}\n\n", add_special_tokens=False) slots, sizes = [], [] many = len(questions) > 1 for k, q in enumerate(questions, 1): opts = ["no", "yes"] if q["kind"] == "noul" else q["options"] tag = {"noul": "yes/no", "score": "score"}.get(q["kind"], "choice") text = f"\nQuestion{' '+str(k) if many else ''} ({tag}): {q['text'].strip()}" text += "\nLevels:" if q["kind"] == "score" else ("\nOptions:" if q["kind"] == "choice" else "") for j, o in enumerate(opts): d = (q.get("descs") or [None] * len(opts))[j] text += f"\n({LETTERS[j]}) {o}" + (f" — {d}" if d else "") text += f"\nAnswer{' '+str(k) if many else ''}: (" ids += tok.encode(text, add_special_tokens=False) slots.append(len(ids) - 1); sizes.append(len(opts)) with torch.no_grad(): logits = model(input_ids=torch.tensor([ids], device=model.device)).logits[0] # [T, vocab] return [torch.softmax(logits[s, LETTER_IDS[:n]].float(), 0).tolist() for s, n in zip(slots, sizes)] state = "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." qs = [ {"kind": "choice", "text": "Which team should handle this ticket?", "options": ["billing", "technical", "account"], "descs": ["payments and refunds", "a product fault", "login or profile settings"]}, {"kind": "noul", "text": "Is the customer asking for a refund?"}, {"kind": "score", "text": "How urgent is this?", "options": ["routine", "soon", "urgent", "critical"]}, ] for q, p in zip(qs, ask(state, qs)): print(q["text"], [round(x, 3) for x in p]) ``` Output for that request: ``` Which team should handle this ticket? [1.0, 0.0, 0.0] # billing Is the customer asking for a refund? [0.0, 1.0] # P(yes) = 1.00 How urgent is this? [0.247, 0.671, 0.08, 0.001] # expected level ≈ 0.8 of 0..3 ``` Rules of the format: yes/no questions always use the options `no`, `yes`; score questions list ordered levels; give option **descriptions** whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template `system_one`. ## On device Use the GGUF repo [`BricksDisplay/jevling-e2b-v1-GGUF`](https://huggingface.co/BricksDisplay/jevling-e2b-v1-GGUF) with the maintained llama.cpp implementation ([`tools/system-one` on mybigday/system-one-llama.cpp, branch `feat/system-one`](https://github.com/mybigday/system-one-llama.cpp/tree/feat/system-one/tools/system-one)). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer. ## Evaluation All numbers are accuracy on datasets the model was **not** trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. | benchmark | task | Jevling-0.8B-v1 | **Jevling-E2B-v1** | |---|---|---|---| | MASSIVE (en-US) | scenario classification, 18-way | 0.667 | **0.742** | | BBC News | topic, 5-way | 0.917 | **0.967** | | TREC | question type, 6-way | 0.758 | **0.792** | | PAWS | paraphrase yes/no | 0.717 | **0.700** | | CommitmentBank | NLI, 3-way | 0.625 | **0.804** | | StrategyQA | yes/no reasoning | 0.500 | **0.525** | | PubMedQA | yes/no/maybe | 0.717 | **0.600** | | SciQ | 4-way science QA | 0.950 | **0.967** | | Social IQa | 3-way | 0.633 | **0.717** | | TruthfulQA (MC) | multiple choice | 0.467 | **0.642** | | XStoryCloze (en) | 2-way | 0.917 | **0.967** | | QuALITY | long-document 4-way QA | 0.425 | **0.567** | | RewardBench | pairwise preference | 0.567 | **0.733** | | Hermes function-calling | tool choice | 0.971 | **0.963** | | Financial PhraseBank | sentiment, 3-way | 0.658 | **0.667** | | JevBench easy / original / hard (231 items) | typed decisions | 1.000 / 0.861 / 0.450 | **1.000 / 0.944 / 0.432** | | zh-TW kiosk set (ours, **synthetic-derived**, 255 states) | intent acc / completeness AUROC / is-order / noise / size | 0.961 / 0.975 / 0.984 / 0.992 / 1.000 | **0.980 / 0.989 / 0.984 / 0.992 / 1.000** | JevBench *hard* (.45) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); the zh-TW kiosk set is our own synthetic-derived data, so read that row as "fit for the distribution it was built for", not as a general claim. ## Limitations - Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.6); compute arithmetic in code and put the result in the state. - Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it. - When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention. - Chat still works: the fine-tune only trains the answer slot, and spot checks show the base model's chat replies are essentially unchanged (chat quality was not benchmarked). The calibration temperature (T = 1.10) is folded into the final norm, so sampled chat output is slightly flatter than the base model's at the same sampling temperature; greedy decoding is unaffected. ## Training data Fine-tuned on commercially licensed public classification / QA / preference / tool-use / safety datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training. ## Licence Apache-2.0 (same as the Gemma-4 base). Trained only on commercially usable data: public datasets under MIT, Apache-2.0, CC-BY-4.0, CC-BY-2.0, CC0, ODC-BY and CDLA-Sharing licences, plus our own synthetic data. CC-BY / ODC-BY sources require attribution; the per-dataset list is available on request.