hans00 commited on
Commit
fcec080
·
verified ·
1 Parent(s): 970aafe

Name the gemma-4 model E2B; note retained chat ability

Browse files
Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -6,16 +6,16 @@ tags: [system-one, jev, typed-decisions, calibrated-classification, kiosk]
6
  library_name: transformers
7
  pipeline_tag: text-classification
8
  ---
9
- # Jevling-2B-v1
10
 
11
- **Jevling-2B-v1** is a small *System One* decision model in the family of TypeSafe's Jev: you give it a **state** (any text — a transcript, a ticket, a document) and one or more **typed questions** (choice / yes-no / score), and it answers all of them in **one forward pass, with no text generation**, each as a calibrated probability distribution over the options. It is fine-tuned from `google/gemma-4-E2B-it` for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).
12
 
13
  ## Quick start (transformers)
14
  ```python
15
  import torch
16
  from transformers import AutoTokenizer, AutoModelForCausalLM
17
 
18
- MODEL = "BricksDisplay/jevling-2b-v1"
19
  tok = AutoTokenizer.from_pretrained(MODEL)
20
  model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval() # on ROCm add attn_implementation="eager"
21
  LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
@@ -61,12 +61,12 @@ How urgent is this? [0.247, 0.671, 0.08, 0.001] # expected level
61
  Rules of the format: yes/no questions always use the options `no`, `yes`; score questions list ordered levels; give option **descriptions** whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template `system_one`.
62
 
63
  ## On device
64
- Use the GGUF repo [`BricksDisplay/jevling-2b-v1-GGUF`](https://huggingface.co/BricksDisplay/jevling-2b-v1-GGUF) with the maintained llama.cpp implementation ([`tools/system-one` on mybigday/system-one-llama.cpp, branch `feat/system-one`](https://github.com/mybigday/system-one-llama.cpp/tree/feat/system-one/tools/system-one)). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.
65
 
66
  ## Evaluation
67
  All numbers are accuracy on datasets the model was **not** trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted.
68
 
69
- | benchmark | task | Jevling-0.8B-v1 | **Jevling-2B-v1** |
70
  |---|---|---|---|
71
  | MASSIVE (en-US) | scenario classification, 18-way | 0.667 | **0.742** |
72
  | BBC News | topic, 5-way | 0.917 | **0.967** |
@@ -92,7 +92,7 @@ JevBench *hard* (.45) is where every open model we know of sits (TypeSafe's Jev
92
  - Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.6); compute arithmetic in code and put the result in the state.
93
  - Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
94
  - When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
95
- - Not a chat model: it does not generate text.
96
 
97
  ## Training data
98
  Fine-tuned on commercially licensed public classification / QA / preference / tool-use / safety datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.
 
6
  library_name: transformers
7
  pipeline_tag: text-classification
8
  ---
9
+ # Jevling-E2B-v1
10
 
11
+ **Jevling-E2B-v1** is a small *System One* decision model in the family of TypeSafe's Jev: you give it a **state** (any text — a transcript, a ticket, a document) and one or more **typed questions** (choice / yes-no / score), and it answers all of them in **one forward pass, with no text generation**, each as a calibrated probability distribution over the options. It is fine-tuned from `google/gemma-4-E2B-it` for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).
12
 
13
  ## Quick start (transformers)
14
  ```python
15
  import torch
16
  from transformers import AutoTokenizer, AutoModelForCausalLM
17
 
18
+ MODEL = "BricksDisplay/jevling-e2b-v1"
19
  tok = AutoTokenizer.from_pretrained(MODEL)
20
  model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval() # on ROCm add attn_implementation="eager"
21
  LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
 
61
  Rules of the format: yes/no questions always use the options `no`, `yes`; score questions list ordered levels; give option **descriptions** whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template `system_one`.
62
 
63
  ## On device
64
+ Use the GGUF repo [`BricksDisplay/jevling-e2b-v1-GGUF`](https://huggingface.co/BricksDisplay/jevling-e2b-v1-GGUF) with the maintained llama.cpp implementation ([`tools/system-one` on mybigday/system-one-llama.cpp, branch `feat/system-one`](https://github.com/mybigday/system-one-llama.cpp/tree/feat/system-one/tools/system-one)). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.
65
 
66
  ## Evaluation
67
  All numbers are accuracy on datasets the model was **not** trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted.
68
 
69
+ | benchmark | task | Jevling-0.8B-v1 | **Jevling-E2B-v1** |
70
  |---|---|---|---|
71
  | MASSIVE (en-US) | scenario classification, 18-way | 0.667 | **0.742** |
72
  | BBC News | topic, 5-way | 0.917 | **0.967** |
 
92
  - Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.6); compute arithmetic in code and put the result in the state.
93
  - Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
94
  - When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
95
+ - Chat still works: the fine-tune only trains the answer slot, and spot checks show the base model's chat replies are essentially unchanged (chat quality was not benchmarked). The calibration temperature (T = 1.10) is folded into the final norm, so sampled chat output is slightly flatter than the base model's at the same sampling temperature; greedy decoding is unaffected.
96
 
97
  ## Training data
98
  Fine-tuned on commercially licensed public classification / QA / preference / tool-use / safety datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.