SanSi

Project page arXiv Code SanSi-2.6B Apache 2.0

SanSi is the model of the paper SanSi: A Looped Typed Decision Model for System 1.5 Thinking. It is a typed decision model: given a text (the state), a question and 2 to 26 declared options, it returns a probability for each option, without generating text. It is built on Ouro-1.4B, a language model pre-trained to loop: one stack of 24 shared layers is applied T = 8 times. The option probabilities are read after every loop, so one forward pass gives the decision at every budget from one loop to eight.

This repository holds what was trained: LoRA adapters of rank 64 and one small readout per loop, 60.8M parameters. The backbone stays frozen; it is downloaded from ByteDance/Ouro-1.4B at the revision the model was trained on.

Model Backbone Trained parameters Accuracy after loop 8
SanSi (this model) Ouro-1.4B 60.8M 72.0
SanSi-2.6B Ouro-2.6B 121.4M 75.8

Accuracy (%) on the 10,027 test items of the paper, mean of three training seeds.

How to use

The loading code is in the GitHub repository. Use a GPU; the backbone runs in bfloat16.

git clone https://github.com/minnesotanlp/Sansi
cd Sansi
pip install -r requirements.txt
from sansi.hub import load, decide

model, tok = load("minnesotanlp/SanSi")   # this adapter, and Ouro-1.4B at its pinned revision

probs = decide(
    model, tok,
    state="Mia is taller than Sam. Sam is taller than Lee.",
    question="Who is the shortest?",
    options=["Mia", "Sam", "Lee"],
)
print(probs[0])    # after loop 1, at an eighth of the computation: about [0.02, 0.07, 0.91]
print(probs[-1])   # after loop 8: about [0.00, 0.00, 1.00]

decide returns one list per loop, each with one probability per option in the order given. It writes the prompt the model was trained on,

<state>

Question: <question>
Options: (A) <option 1> (B) <option 2> ...
Answer:

and reads the option letters at its last token after every loop. decide(..., loops=4) runs four loops only, at half the computation (71.6% instead of 72.0% on the test items). To run a whole dataset, use eval/evaluate.py of the repository. An item whose answer the state does not give was trained towards the uniform distribution: a low top probability means that the model does not commit to an answer (the paper counts a hard answer when the top probability reaches (1 + 1/K) / 2 for K options).

Results

Accuracy (%) on the 10,027 test items after every loop: the mean of the three training seeds of the paper, and this checkpoint (seed 0).

Loop 1 2 3 4 5 6 7 8
Mean of three seeds 58.4 66.9 70.4 71.6 71.9 72.1 72.1 72.0
This checkpoint (seed 0) 57.7 67.2 70.4 71.2 71.2 71.4 71.4 71.3

The main comparison of the paper: all models are trained on the same data, with the same recipe except Kev-4B (our data), which uses Kev's code; mean of three seeds. In-distribution items come from the training sources, near transfer from harder or reworded versions of them, far transfer from sources never trained on. Cost is the GPU time of one pass over the test items, with one loop of Ouro-1.4B as 1.

Model Params Loops Cost All In-dist. Near Far ECE
SmolLM2-1.7B 1.7B 1 1.1 58.4 76.1 57.7 52.3 0.069
Ouro-1.4B, one loop 1.4B 1 1.0 58.6 77.3 59.9 51.3 0.137
Qwen3.5-2B 1.9B 1 1.2 66.7 84.3 62.3 61.7 0.123
Qwen3.5-4B 4.2B 1 2.4 73.8 88.5 68.9 70.1 0.113
Kev-4B (our data) 4.2B 1 2.1 74.3 88.6 70.9 70.3 0.119
SanSi 1.4B 8 7.7 72.0 86.6 67.7 68.0 0.093
SanSi-2.6B 2.7B 8 14.8 75.8 88.4 74.4 71.6 0.078

Training

  • Data: the 12,800 training items of the paper's decision suite, from 20 public sources. The GitHub repository downloads the sources at fixed revisions and rebuilds the dataset byte for byte (data.download_raw, data.build_main, data.verify).
  • Model: Ouro-1.4B frozen; LoRA of rank 64 (alpha 128, dropout 0.05) on the attention and MLP projections; one readout of rank 16 per loop.
  • Loss at every loop: the cross-entropy plus the Brier score against the item's target distribution.
  • 1,000 steps of 16 items; AdamW with learning rates 1e-4 (LoRA) and 1e-3 (readout), 200 warm-up steps, cosine decay to 10%, gradient clipping at 1.0; two RTX A6000 GPUs; seed 0.
python -m train.train --model <Ouro-1.4B> --loops 8 --seed 0 --out runs/sansi_s0

Limitations

  • English only. The model answers one question with 2 to 26 declared options per call, in the prompt format above.
  • Running more loops than the eight it was trained with lowers accuracy.
  • This is one training seed; the test accuracy of the three seeds of the paper has a standard deviation of 0.7 points.
  • Looping does not replace parameters entirely: Qwen3.5-4B, with three times the parameters, is 1.8 points more accurate, at a third of the computation.

License

The adapter and the readouts are released under the Apache 2.0 license, as are the code and the backbone Ouro-1.4B. The training data comes from public datasets that keep their own licenses (data/README.md of the repository).

Citation

@misc{gan2026sansiloopedtypeddecision,
      title={SanSi: A Looped Typed Decision Model for System 1.5 Thinking},
      author={Shuyu Gan and Young-Jun Lee and Dongyeop Kang},
      year={2026},
      eprint={2610.07730},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.07730},
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for minnesotanlp/SanSi

Adapter
(2)
this model

Collection including minnesotanlp/SanSi

Paper for minnesotanlp/SanSi