File size: 7,358 Bytes
e836125
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b68d425
e836125
 
b68d425
e836125
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b68d425
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
---

language:
- en
- zh
license: apache-2.0
tags:
- decision-model
- calibration
- system-one
- from-scratch
- probabilistic-decisions
- pytorch
pretty_name: MiniSystemOne
---


# MiniSystemOne

**Train a probabilistic decision model from scratch β€” no LLM, no decoding, no JSON generation.**

A 26.89M-parameter model that takes a **state** plus **typed questions** and returns
**typed decisions with calibrated probabilities** β€” in a single parallel forward pass.
No autoregressive loop, no text output, no constrained-decoding tricks.

Built from **random initialization**, MiniMind-style: one shared bidirectional encoder,
one decision head, one loss. Trained in **4.7 hours** on a single RTX 4070 Laptop (8 GB).

> ### ⚠️ This is not a Jev reproduction
>

> It is an **independent educational implementation**, inspired by the "System One Model"
> idea popularized by TypeSafe AI's Jev (2026-09-15) β€” the way MiniMind is an educational
> reimplementation of the LLaMA/GPT recipe rather than a reproduction of any model.
> No Jev output was used in training, and no attempt was made to reverse-engineer it.
> It does **not** claim to beat Jev on any axis.

---

## Files

| File | Stage | Params | What it is |
|---|---|---|---|
| `decision.pth` | 2 | 26.89M | **The decision model.** The artifact this repo is about. |
| `mlm.pth` | 1 | 25.31M | The MLM-pretrained encoder. `decision.pth` was warm-started from it. |
| `tokenizer.json` / `tokenizer_config.json` | β€” | vocab 6400 | A **new** BPE, 11 special tokens. Not MiniMind's. |

Both checkpoints are **self-describing**: the file carries the config it was built from
and the provenance of its training data.

```

decision.pth  format=1  stage=decision  step=33795  n_params=26889729

              tokenizer_sha1=bf7a131a109ea445  gen_version=1.0.0

              encoder_init=mlm.pth  trained_on=synth  max_len=1024  epochs=3



mlm.pth       format=1  stage=mlm  step=66000  n_params=25309184

              tokenizer_sha1=bf7a131a109ea445  synth_gen_version=1.0.0

              max_len=512  epochs=8

```

---

## How to use these weights

**There is no `from_pretrained` path, on purpose.** This is a decision-native architecture

with a custom input packing format, not a causal LM. The packing is canonical and lives in

exactly one place in the source repository; re-implementing it here would create a second

implementation that drifts from the first.



The code is at **`github.com/Colvin0315/MiniSystemOne`** (Apache-2.0).



```bash

git clone https://github.com/Colvin0315/MiniSystemOne && cd MiniSystemOne

pip install -r requirements.txt          # torch 2.6.0+cu124, transformers, datasets



# put decision.pth + tokenizer.json + tokenizer_config.json into model/ and out/decision/

python eval/eval_harness.py --ckpt out/decision/decision.pth \

    --data dataset/synth --sets test_known

```



Loading a checkpoint is one call, and it reads the embedded config and refuses a mismatch:



```python

from trainer.trainer_utils import ckpt_info, init_model, verify_tokenizer

from model.model_system_one import DecisionConfig, MiniSystemOneForDecision



verify_tokenizer("decision.pth", "model")        # rejects a mismatched vocabulary

cfg = DecisionConfig(**{k: v for k, v in ckpt_info("decision.pth")["config"].items()

                        if k in ("hidden_size", "num_hidden_layers", "vocab_size",

                                 "num_attention_heads", "num_key_value_heads",

                                 "intermediate_size")})

model = init_model(MiniSystemOneForDecision, cfg, "decision.pth", "cuda")

```



---



## Results



All numbers are from `decision.pth`, reproducible from the GitHub repo.



### Calibration β€” the point of the whole project



One-hot labels **cannot** teach calibration: a Brier loss on one-hot targets is just a

confidence-pushing regularizer. The training data therefore contains samples whose **true

conditional distribution is known** β€” a known randomized rule (`explicit_rng`),

marginalization over a hidden variable (`marginalized`), genuinely tied answer sets

(`tie_set`), and real human disagreement from ChaosNLI (`human_annotators`).



| Synthetic `test_known` (n=18,000, K≀255) | acc | ECE | Brier |

|---|---|---|---|

| uncalibrated | 0.647 | 0.0047 | 0.0249 |

| after global temperature | 0.647 | 0.0038 | 0.0249 |

| **calibration subset** (one-hot `hard` excluded, n=16,361) | **0.612** | **0.0061** | 0.0273 |



**Temperature calibration is a no-op** β€” fitted `T = 0.965`, NLL moves 1.0629 β†’ 1.0628.

That is the intended result: the model is *natively* calibrated because its targets were

distributions to begin with.



### Read this table, not the accuracy column



`accuracy` pools two kinds of sample with structurally different ceilings, so it measures

the provenance mix more than the model. On a `tie_set` record the target is uniform over

the valid answers, so soft-accuracy is capped at `1/k` no matter how good the model is.

The honest denominator is the **oracle ceiling** `mean(max_k t_k)` β€” computable from the
data with no model at all.

| generator | model | oracle ceiling | achieved |
|---|---|---|---|
| `tool_router` | 0.7137 | 0.7140 | **100.0%** |
| `security_gate` | 0.6403 | 0.6426 | 99.6% |
| `refund_policy` | 0.7561 | 0.7601 | 99.5% |
| `agent_trace_score` | 0.6366 | 0.6450 | 98.7% |
| `banking_balance` | 0.4544 | 0.6091 | 74.6% |
| `calendar_slot` | 0.0920 | 0.1416 | 65.0% |

Four of six are essentially solved. `banking_balance` is the arithmetic one and is also
where calibration is worst (ECE 0.1637 vs 0.0035–0.0495 elsewhere); `calendar_slot` is the
large-K one, ambiguous by construction.

### Where it fails β€” real text

| ChaosNLI (474 items, Nβ‰ˆ100 annotators) | |
|---|---|
| ECE | **0.0613** |
| binomial noise floor | 0.0068 |
| noise-corrected | 0.0545 |

The floor explains only 11% of it. **On real human disagreement the model is genuinely

miscalibrated**, and no noise correction rescues that claim.

**Synthetic β†’ real gap: 0.612 β†’ 0.424 soft-accuracy β€” about 19 points lost** crossing from
program-generated rules to real natural language. This is the most informative single
number this project produces, reported as a limitation rather than a footnote.

### Efficiency

| | |
|---|---|
| latency (B=1, per request) | 20.24 ms median, 27.09 ms p95 |
| peak VRAM | **0.13 GB**, flat across K ∈ {2, 32, 128, 255} |
| K=255 | 49.97 ms median (chunked path) |

Latency is **overhead-bound, not compute-bound**: K=2 takes 19.88 ms and K=32 takes
18.37 ms β€” 30 extra candidates cost nothing.

---

## What it cannot do

- **Match an LLM's open-domain NLU.** It is a **schema-bound decision model**, not a
  general assistant.
- **Win on accuracy against a language model.** The claims are latency, native calibrated
  distributions, and **zero schema errors by construction** β€” not accuracy.
- **Give reliable probabilities on genuinely OOD inputs.** No model does. The ChaosNLI
  figure above shows the degradation rather than claiming otherwise.
- **Replace an LLM in any sense.** It is a component.

## License

Apache-2.0 β€” see the [source repository](https://github.com/Colvin0315/MiniSystemOne).