File size: 6,299 Bytes
8e7b1d3 eef2655 8e7b1d3 eef2655 8e7b1d3 eef2655 be55743 eef2655 b49b31c eef2655 b49b31c eef2655 b49b31c eef2655 be55743 8e7b1d3 eef2655 8e7b1d3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | ---
license: apache-2.0
base_model: google/gemma-4-E2B-it
tags:
- decision-model
- system-1
- rlcd
- proper-scoring-rules
- gemma
- classone
- classification
pipeline_tag: text-classification
---
# classone-gemma4-e2b — ClassOne System 1 Decision Model
**[devops-thiago/classone-gemma4-e2b](https://huggingface.co/devops-thiago/classone-gemma4-e2b)** is an open-source **System 1 decision model** using the [ClassOne architecture](https://github.com/devops-thiago/class-one). The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download.
Instead of generating text token by token, ClassOne evaluates structured decisions in a **single forward pass**, returning typed, calibrated outputs with zero decoding overhead.
## Benchmark Results
### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks)
Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench):
| Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) |
|---|---|---|---|---|---|
| **Easy** | 48 | **93.8%** (45/48) | 0.0610 | 0.0516 | **45.0 ms** |
| **Original** | 72 | **63.9%** (46/72) | 0.2510 | 0.2499 | **43.1 ms** |
| **Hard** | 111 | **37.8%** (42/111) | 0.4299 | 0.4022 | **91.3 ms** |
| **Overall Aggregate** | **231** | **57.6%** (133/231) | — | — | **~44 ms** |
- **Easy Tier Sub-Breakdown:** Choice accuracy: **100.0%** (36/36); Noul policy accuracy: **75.0%** (9/12).
- **Original Tier Sub-Breakdown:** Choice accuracy: **66.7%** (24/36); Score rubrics: **66.7%** (8/12); Noul accuracy: **58.3%** (14/24).
- **Hard Tier Sub-Breakdown:** Noul policy compliance: **44.7%** (17/38); Choice accuracy: **34.3%** (23/67); Score rubrics: **33.3%** (2/6).
### 2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances)
Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429):
| Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) |
|---|---|---|---|---|---|
| **Concealing Uncertainty** | 14 | **0.980** | **85.7%** | **0.0718** | 130.2 ms |
| **Honesty (Deception)** | 11 | **0.800** | **72.7%** | 0.2445 | 219.6 ms |
| **Refusal (Jailbreaks)** | 11 | **0.667** | **54.5%** | 0.1934 | 271.1 ms |
| **Power Seeking** | 6 | **0.556** | 50.0% | 0.1794 | 213.2 ms |
| **Reward Hacking** | 9 | **0.500** | 33.3% | 0.2935 | 209.6 ms |
| **Prompt Injection** | 8 | **0.500** | 37.5% | 0.3207 | 167.9 ms |
| **Bias** | 9 | 0.375 | 55.6% | 0.1659 | 221.5 ms |
| **Overall Average** | **100** | **0.516** | **51.0%** | **0.1584** | **200.0 ms** |
### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API)
Measured against TypeSafe AI's Jev (v1.13) cloud API:
- **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private)
- **TypeSafe Jev (Cloud API):** **329.90 ms** mean latency (3.0 req/s)
- **Edge Speedup:** **6.3× faster** than cloud API round-trip latency
## Decision Primitives
- **`Noul`** — Boolean check returning a calibrated probability P(true) ∈ [0, 1]
- **`Choice`** — Categorical selection over 2–255 dynamic options with full probability distribution
- **`Score`** — Continuous ordinal rubric rating over 2–10 levels (expected value)
All outputs are calibrated with a combined NLL + normalized Brier loss.
Post-hoc temperature calibration achieves **ECE = 0.034** (down from 0.178).
## Quickstart
```bash
pip install classone
```
```python
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
from classone.modeling.modeling_classone import ClassOneModel
from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion
from classone.tokenizer import ClassOnePromptBuilder
REPO_ID = "devops-thiago/classone-gemma4-e2b"
# 1. Load the ClassOne model (weights + tokenizer are fully self-contained here)
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
builder = ClassOnePromptBuilder(tokenizer)
model = ClassOneModel.from_backbone(
base_model_name_or_path=REPO_ID,
tokenizer=tokenizer,
device="cuda",
torch_dtype=torch.float16,
)
# 2. Load the trained decision heads
heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda")
model.noul_head.load_state_dict(heads["noul_head"])
model.choice_head.load_state_dict(heads["choice_head"])
model.score_head.load_state_dict(heads["score_head"])
model.eval()
# 3. Pack state + questions and run a single forward pass
packed = builder.pack(
state={"customer": "Alex", "message": "I was charged twice for order #123."},
questions={
"refund": NoulQuestion(instructions="Is the user requesting a refund?"),
"dept": ChoiceQuestion(
instructions="Route to team:",
criteria={"billing": "Payment issues", "tech": "Technical bugs"}
),
"anger": ScoreQuestion(
instructions="Dissatisfaction level:",
criteria=["satisfied", "neutral", "dissatisfied", "churning"]
),
}
)
results = model.evaluate_packed(packed)
print("Refund P(true):", results["refund"].noul)
print("Department: ", results["dept"].choice, "—", results["dept"].probabilities)
print("Anger score: ", results["anger"].score)
```
## Repository Files
| File | Description |
|---|---|
| `model.safetensors` (sharded) | Merged ClassOne backbone weights |
| `config.json` | Model configuration |
| `tokenizer.json`, `tokenizer_config.json` | Tokenizer, including ClassOne delimiter tokens |
| `classone_heads.pt` | Trained Noul / Choice / Score head weights + calibrated temperatures |
| `lora_backbone/` | LoRA adapter (r=16, α=32) that produced the merged weights |
## Citation
```bibtex
@misc{classone2026,
title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models},
author={Thiago Gonzaga},
year={2026},
url={https://github.com/devops-thiago/class-one},
}
```
## Attribution & Legal
- Derived from [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (Google) — Apache License 2.0
- Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0
|