Instructions to use minnesotanlp/SanSi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use minnesotanlp/SanSi with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("ByteDance/Ouro-1.4B") model = PeftModel.from_pretrained(base_model, "minnesotanlp/SanSi") - Notebooks
- Google Colab
- Kaggle
SanSi
SanSi is the model of the paper SanSi: A Looped Typed Decision Model for System 1.5 Thinking. It is a typed decision model: given a text (the state), a question and 2 to 26 declared options, it returns a probability for each option, without generating text. It is built on Ouro-1.4B, a language model pre-trained to loop: one stack of 24 shared layers is applied T = 8 times. The option probabilities are read after every loop, so one forward pass gives the decision at every budget from one loop to eight.
This repository holds what was trained: LoRA adapters of rank 64 and one small readout per loop, 60.8M parameters.
The backbone stays frozen; it is downloaded from ByteDance/Ouro-1.4B at the revision the model was trained on.
| Model | Backbone | Trained parameters | Accuracy after loop 8 |
|---|---|---|---|
| SanSi (this model) | Ouro-1.4B | 60.8M | 72.0 |
| SanSi-2.6B | Ouro-2.6B | 121.4M | 75.8 |
Accuracy (%) on the 10,027 test items of the paper, mean of three training seeds.
How to use
The loading code is in the GitHub repository. Use a GPU; the backbone runs in bfloat16.
git clone https://github.com/minnesotanlp/Sansi
cd Sansi
pip install -r requirements.txt
from sansi.hub import load, decide
model, tok = load("minnesotanlp/SanSi") # this adapter, and Ouro-1.4B at its pinned revision
probs = decide(
model, tok,
state="Mia is taller than Sam. Sam is taller than Lee.",
question="Who is the shortest?",
options=["Mia", "Sam", "Lee"],
)
print(probs[0]) # after loop 1, at an eighth of the computation: about [0.02, 0.07, 0.91]
print(probs[-1]) # after loop 8: about [0.00, 0.00, 1.00]
decide returns one list per loop, each with one probability per option in the order given. It writes the prompt
the model was trained on,
<state>
Question: <question>
Options: (A) <option 1> (B) <option 2> ...
Answer:
and reads the option letters at its last token after every loop. decide(..., loops=4) runs four loops only, at half
the computation (71.6% instead of 72.0% on the test items). To run a whole dataset, use eval/evaluate.py of the
repository. An item whose answer the state does not give was trained towards the uniform distribution: a low top
probability means that the model does not commit to an answer (the paper counts a hard answer when the top
probability reaches (1 + 1/K) / 2 for K options).
Results
Accuracy (%) on the 10,027 test items after every loop: the mean of the three training seeds of the paper, and this checkpoint (seed 0).
| Loop | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Mean of three seeds | 58.4 | 66.9 | 70.4 | 71.6 | 71.9 | 72.1 | 72.1 | 72.0 |
| This checkpoint (seed 0) | 57.7 | 67.2 | 70.4 | 71.2 | 71.2 | 71.4 | 71.4 | 71.3 |
The main comparison of the paper: all models are trained on the same data, with the same recipe except Kev-4B (our data), which uses Kev's code; mean of three seeds. In-distribution items come from the training sources, near transfer from harder or reworded versions of them, far transfer from sources never trained on. Cost is the GPU time of one pass over the test items, with one loop of Ouro-1.4B as 1.
| Model | Params | Loops | Cost | All | In-dist. | Near | Far | ECE |
|---|---|---|---|---|---|---|---|---|
| SmolLM2-1.7B | 1.7B | 1 | 1.1 | 58.4 | 76.1 | 57.7 | 52.3 | 0.069 |
| Ouro-1.4B, one loop | 1.4B | 1 | 1.0 | 58.6 | 77.3 | 59.9 | 51.3 | 0.137 |
| Qwen3.5-2B | 1.9B | 1 | 1.2 | 66.7 | 84.3 | 62.3 | 61.7 | 0.123 |
| Qwen3.5-4B | 4.2B | 1 | 2.4 | 73.8 | 88.5 | 68.9 | 70.1 | 0.113 |
| Kev-4B (our data) | 4.2B | 1 | 2.1 | 74.3 | 88.6 | 70.9 | 70.3 | 0.119 |
| SanSi | 1.4B | 8 | 7.7 | 72.0 | 86.6 | 67.7 | 68.0 | 0.093 |
| SanSi-2.6B | 2.7B | 8 | 14.8 | 75.8 | 88.4 | 74.4 | 71.6 | 0.078 |
Training
- Data: the 12,800 training items of the paper's decision suite, from 20 public sources. The GitHub repository
downloads the sources at fixed revisions and rebuilds the dataset byte for byte (
data.download_raw,data.build_main,data.verify). - Model: Ouro-1.4B frozen; LoRA of rank 64 (alpha 128, dropout 0.05) on the attention and MLP projections; one readout of rank 16 per loop.
- Loss at every loop: the cross-entropy plus the Brier score against the item's target distribution.
- 1,000 steps of 16 items; AdamW with learning rates 1e-4 (LoRA) and 1e-3 (readout), 200 warm-up steps, cosine decay to 10%, gradient clipping at 1.0; two RTX A6000 GPUs; seed 0.
python -m train.train --model <Ouro-1.4B> --loops 8 --seed 0 --out runs/sansi_s0
Limitations
- English only. The model answers one question with 2 to 26 declared options per call, in the prompt format above.
- Running more loops than the eight it was trained with lowers accuracy.
- This is one training seed; the test accuracy of the three seeds of the paper has a standard deviation of 0.7 points.
- Looping does not replace parameters entirely: Qwen3.5-4B, with three times the parameters, is 1.8 points more accurate, at a third of the computation.
License
The adapter and the readouts are released under the Apache 2.0 license, as are the code and the backbone Ouro-1.4B.
The training data comes from public datasets that keep their own licenses (data/README.md of the repository).
Citation
@misc{gan2026sansiloopedtypeddecision,
title={SanSi: A Looped Typed Decision Model for System 1.5 Thinking},
author={Shuyu Gan and Young-Jun Lee and Dongyeop Kang},
year={2026},
eprint={2610.07730},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2610.07730},
}
- Downloads last month
- -
Model tree for minnesotanlp/SanSi
Base model
ByteDance/Ouro-1.4B