| --- |
| license: apache-2.0 |
| base_model: google/gemma-4-E2B-it |
| tags: |
| - decision-model |
| - system-1 |
| - rlcd |
| - proper-scoring-rules |
| - gemma |
| - classone |
| - classification |
| pipeline_tag: text-classification |
| --- |
| |
| # classone-gemma4-e2b — ClassOne System 1 Decision Model |
|
|
| **[devops-thiago/classone-gemma4-e2b](https://huggingface.co/devops-thiago/classone-gemma4-e2b)** is an open-source **System 1 decision model** using the [ClassOne architecture](https://github.com/devops-thiago/class-one). The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download. |
|
|
| Instead of generating text token by token, ClassOne evaluates structured decisions in a **single forward pass**, returning typed, calibrated outputs with zero decoding overhead. |
|
|
| ## Benchmark Results |
|
|
| ### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks) |
|
|
| Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench): |
|
|
| | Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) | |
| |---|---|---|---|---|---| |
| | **Easy** | 48 | **93.8%** (45/48) | 0.0610 | 0.0516 | **45.0 ms** | |
| | **Original** | 72 | **63.9%** (46/72) | 0.2510 | 0.2499 | **43.1 ms** | |
| | **Hard** | 111 | **37.8%** (42/111) | 0.4299 | 0.4022 | **91.3 ms** | |
| | **Overall Aggregate** | **231** | **57.6%** (133/231) | — | — | **~44 ms** | |
|
|
| - **Easy Tier Sub-Breakdown:** Choice accuracy: **100.0%** (36/36); Noul policy accuracy: **75.0%** (9/12). |
| - **Original Tier Sub-Breakdown:** Choice accuracy: **66.7%** (24/36); Score rubrics: **66.7%** (8/12); Noul accuracy: **58.3%** (14/24). |
| - **Hard Tier Sub-Breakdown:** Noul policy compliance: **44.7%** (17/38); Choice accuracy: **34.3%** (23/67); Score rubrics: **33.3%** (2/6). |
|
|
| ### 2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances) |
|
|
| Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429): |
|
|
| | Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) | |
| |---|---|---|---|---|---| |
| | **Concealing Uncertainty** | 14 | **0.980** | **85.7%** | **0.0718** | 130.2 ms | |
| | **Honesty (Deception)** | 11 | **0.800** | **72.7%** | 0.2445 | 219.6 ms | |
| | **Refusal (Jailbreaks)** | 11 | **0.667** | **54.5%** | 0.1934 | 271.1 ms | |
| | **Power Seeking** | 6 | **0.556** | 50.0% | 0.1794 | 213.2 ms | |
| | **Reward Hacking** | 9 | **0.500** | 33.3% | 0.2935 | 209.6 ms | |
| | **Prompt Injection** | 8 | **0.500** | 37.5% | 0.3207 | 167.9 ms | |
| | **Bias** | 9 | 0.375 | 55.6% | 0.1659 | 221.5 ms | |
| | **Overall Average** | **100** | **0.516** | **51.0%** | **0.1584** | **200.0 ms** | |
|
|
| ### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API) |
|
|
| Measured against TypeSafe AI's Jev (v1.13) cloud API: |
| - **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private) |
| - **TypeSafe Jev (Cloud API):** **329.90 ms** mean latency (3.0 req/s) |
| - **Edge Speedup:** **6.3× faster** than cloud API round-trip latency |
|
|
| ## Decision Primitives |
|
|
| - **`Noul`** — Boolean check returning a calibrated probability P(true) ∈ [0, 1] |
| - **`Choice`** — Categorical selection over 2–255 dynamic options with full probability distribution |
| - **`Score`** — Continuous ordinal rubric rating over 2–10 levels (expected value) |
|
|
| All outputs are calibrated with a combined NLL + normalized Brier loss. |
| Post-hoc temperature calibration achieves **ECE = 0.034** (down from 0.178). |
|
|
| ## Quickstart |
|
|
| ```bash |
| pip install classone |
| ``` |
|
|
| ```python |
| import torch |
| from huggingface_hub import hf_hub_download |
| from transformers import AutoTokenizer |
| |
| from classone.modeling.modeling_classone import ClassOneModel |
| from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion |
| from classone.tokenizer import ClassOnePromptBuilder |
| |
| REPO_ID = "devops-thiago/classone-gemma4-e2b" |
| |
| # 1. Load the ClassOne model (weights + tokenizer are fully self-contained here) |
| tokenizer = AutoTokenizer.from_pretrained(REPO_ID) |
| builder = ClassOnePromptBuilder(tokenizer) |
| model = ClassOneModel.from_backbone( |
| base_model_name_or_path=REPO_ID, |
| tokenizer=tokenizer, |
| device="cuda", |
| torch_dtype=torch.float16, |
| ) |
| |
| # 2. Load the trained decision heads |
| heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda") |
| model.noul_head.load_state_dict(heads["noul_head"]) |
| model.choice_head.load_state_dict(heads["choice_head"]) |
| model.score_head.load_state_dict(heads["score_head"]) |
| model.eval() |
| |
| # 3. Pack state + questions and run a single forward pass |
| packed = builder.pack( |
| state={"customer": "Alex", "message": "I was charged twice for order #123."}, |
| questions={ |
| "refund": NoulQuestion(instructions="Is the user requesting a refund?"), |
| "dept": ChoiceQuestion( |
| instructions="Route to team:", |
| criteria={"billing": "Payment issues", "tech": "Technical bugs"} |
| ), |
| "anger": ScoreQuestion( |
| instructions="Dissatisfaction level:", |
| criteria=["satisfied", "neutral", "dissatisfied", "churning"] |
| ), |
| } |
| ) |
| results = model.evaluate_packed(packed) |
| |
| print("Refund P(true):", results["refund"].noul) |
| print("Department: ", results["dept"].choice, "—", results["dept"].probabilities) |
| print("Anger score: ", results["anger"].score) |
| ``` |
|
|
| ## Repository Files |
|
|
| | File | Description | |
| |---|---| |
| | `model.safetensors` (sharded) | Merged ClassOne backbone weights | |
| | `config.json` | Model configuration | |
| | `tokenizer.json`, `tokenizer_config.json` | Tokenizer, including ClassOne delimiter tokens | |
| | `classone_heads.pt` | Trained Noul / Choice / Score head weights + calibrated temperatures | |
| | `lora_backbone/` | LoRA adapter (r=16, α=32) that produced the merged weights | |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{classone2026, |
| title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models}, |
| author={Thiago Gonzaga}, |
| year={2026}, |
| url={https://github.com/devops-thiago/class-one}, |
| } |
| ``` |
|
|
| ## Attribution & Legal |
|
|
| - Derived from [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (Google) — Apache License 2.0 |
| - Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0 |
|
|