- Ines-1
- Supported decision interface
- Intended use
- Quickstart
- Lineage
- Results
- typed-decisions — single-question specialist protocol
- typed-decisions — full-context stress tests
- What else the public tests measure
- External figures, kept apart (no leaderboard ranking)
- External zero-shot classification: BTZSC-22
- Private domain benchmark
- Ablations (public tests, same questions, exact McNemar)
- Speed
- Training
- Training data
- Limitations
- Reproducibility
- Future work (hypotheses, not results)
- Files
- Paper
- Citation
- Supported decision interface
Ines-1
Licences — what covers what.
- Apache-2.0 (LICENSE, NOTICE): the model weights (
model.safetensors),config.jsonandtokenizer/; the documentation (README.mdand the other.mdfiles,paper/,tasksource_license_audit.csv); the training records (training/*.json); the model's own evaluation outputs ineval/;SHA256SUMS.- MIT (LICENSE-CODE): the code —
mini_v41/,mini_v41_jev/,scripts/,examples/,training/code/,eval/btzsc22/code/— and the environment and build files:Dockerfile,.dockerignore,requirements.txt,requirements.lock,.gitattributes,.gitignore.- Third-party terms, not relicensed: the gold labels and teacher distributions of the public test sets that
eval/items/,eval/reference_rows.jsonandeval/btzsc22/(BTZSC label texts and targets) include keep their datasets' licences. The training data is not redistributed and remains under its own terms, including attribution, share-alike and copyleft ones: see THIRD_PARTY_DATA_NOTICE.md and LICENSING_NOTES.md.
Ines-1 is a typed-decision model pretrained from scratch: a 1.59B-parameter mixture-of-experts causal
encoder–decoder (405M parameters active per token) with hashed n-gram memory (Engram), pretrained on about 10B
public tokens on one H200, then fine-tuned to answer one typed question at a time — choice, score or noul
(yes/no) — with a probability for every option, in English and Spanish. The request shape follows the typed-decision
POST /v1/systemone interface popularised by TypeSafe's Jev. Ines-1 does not reproduce Jev's architecture or its
full-case (multi-question) protocol.
Development codename: mini-v41 / mini-v41-Decisions. The Python packages keep that name (mini_v41,
mini_v41_jev), and so do the provenance records in config.json and training/.
Supported decision interface
- One typed question per model prompt. The answer is the softmax over the option letters (at most 26) at the first reply position.
- Several questions of a case can be API-batched: the server (and
Decider.decide) splits a case into one prompt per question and scores them together. This is batching, not multi-question modelling. - The model was not trained to condition jointly on several typed questions in one prompt, and it does not do so well: with the whole case in the prompt its accuracy falls from 76.4 % to 37–42 % (Full-context stress tests).
Intended use
- Typed single-question decisions over a piece of state (text or JSON): classification, routing and scoring over a closed, caller-defined option set, with a probability per option.
- A starting checkpoint for further domain specialisation (fine-tuning on your own decisions).
Out of scope, or not supported by our evidence: use as a chat or general assistant; retrieval or fact lookup in long documents; native multi-question (Jev full-case) compatibility; strong zero-shot performance on a new domain; any safety-critical decision.
| Parameters | 1,590,066,240 total; 405,052,224 active per token; weights in bf16 |
| Layers | 8 encoder + 8 decoder in one causal stack; hidden 1024; 16 heads / 4 KV heads, head dim 64 |
| Decoder | sliding window of 128 + one global K/V projected from the encoder output; receptive field 1,017 positions |
| MoE | every layer: 16 routed experts (FFN 1536), top-2, + 1 shared expert |
| Engram | hashed 2/3/4-gram embeddings at encoder layers 1 and 5 (2 × 10⁶ rows, 128M parameters) |
| Context | 2,048 tokens; longer cases are trimmed in the middle |
| Tokenizer | byte-level BPE, 65,536 tokens (identity sha256 859e5c63…74469, checked at load) |
Quickstart
pip install -r requirements.txt # requirements.lock pins the canonical environment (ENVIRONMENT.md)
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Endikavi/Ines-1") # the whole repository: weights, tokenizer and reader code
sys.path.insert(0, path)
from mini_v41_jev import Decider
decider = Decider(path) # CUDA if available (bf16 + autocast), otherwise CPU (fp32)
question = {"type": "choice",
"instructions": "Which department should handle this ticket?",
"criteria": {"billing": "payments and invoices", "tech": "outages and bugs", "legal": "disputes"}}
state = "The invoice was charged twice and the customer wants the money back today."
answer = decider.decide(state, {"dept": question}, lang="en")["dept"]
print(answer["choice"]) # argmax option
print(answer["probabilities"]) # probability of every option
score questions return probabilities over levels "0".."n-1"; noul questions return {"noul": P(true)}.
lang="es" uses Spanish prompts. examples/quickstart.py runs the same from a local copy.
HTTP server with the same request shape (POST /v1/systemone, GET /v1/models, GET /health; CUDA and aiohttp),
batching concurrent questions into one prefill: python scripts/serve_jev.py --port 30040.
Lineage
Pretrain-V3 (9.96B tokens) → Decision-FT Stage 1 (typed-decisions + jev-decisions; 8 epochs scheduled, 7
effective: the epoch chosen on validation is the checkpoint kept) → Decision-Mix Stage 2 (balanced public mix,
2 epochs) → Ines-1. Each stage's parent-weights hash matches the previous stage's output (table under Training).
Results
All numbers measured by us on the released weights in the canonical environment (ENVIRONMENT.md:
bf16 weights under bf16 autocast). Per-item outputs in eval/items/, reports in eval/. Accuracy = the
highest-probability option equals the gold label; intervals are bootstrap over cases (the five questions of a case
are not independent).
typed-decisions — single-question specialist protocol
Single-question specialist protocol, after training on the benchmark's training split: one prompt per question;
the model was trained on typed-decisions train (same four workflows, with the dataset's teacher distributions).
This is not a zero-shot result.
| test (400 cases) | correct / n | accuracy | 95 % CI | choice | score | noul | Brier | KL(gold‖p) | log loss |
|---|---|---|---|---|---|---|---|---|---|
| English | 1528 / 2000 | 76.4 | 74.2–78.5 | 71.5 | 74.5 | 83.8 | 0.065 | 0.121 | 0.635 |
| Spanish (our machine translation) | 1496 / 1985 | 75.4 | 73.2–77.4 | 69.9 | 73.9 | 82.7 | 0.065 | 0.119 | 0.639 |
| Spanish (independent translation) | 1526 / 2000 | 76.3 | 74.2–78.5 | 73.0 | 73.8 | 83.0 | 0.064 | 0.119 | 0.639 |
Brier and KL are against the teacher's gold distribution and follow the dataset card's definitions (they reproduce
its Uniform reference row exactly: KL 0.444, Brier 0.238). Our ECE (10 equal-width bins of top-1 confidence; 0.132
on English) is reported as our-ECE and is not comparable with published ECE values, whose code is not public.
By teacher agreement (label_agreement.argmax_agree): 86.4 % where the teacher's samples agree (n = 1,188),
61.7 % where they do not (n = 812).
typed-decisions — full-context stress tests
The dataset's public runs send the whole case (state + its five questions) in one request, and its card notes that request shape can change results. Ines-1 has no readout that answers several questions in one forward, so we ran two stress tests instead (English; Spanish behaves the same):
| readout | what the model sees and how answers are read | correct / n | accuracy | 95 % CI | Brier |
|---|---|---|---|---|---|
| single question (above) | the state and one question | 1528 / 2000 | 76.4 | 74.2–78.5 | 0.065 |
context |
the state and all five questions, ending "Answer question k"; one forward per question | 831 / 2000 | 41.5 | 39.5–43.6 | 0.266 |
joint |
one prompt with all five questions; answers read in sequence as "k: letter", each conditioned on the model's previous answers | 742 / 2000 | 37.1 | 34.9–39.3 | 0.246 |
| dataset reference: Prior | train label frequencies, ignoring the input | 47.0 | 0.189 |
Neither stress test reproduces the public protocol or a Jev-style single forward over five independent decisions;
neither is a DecisionEval score. Both show strong sensitivity to request shape. Decision-FT Stage 1 drops the same
way (75.5 → 42.8 %, context), so this is not caused by Stage 2. Task competence does not imply interface
generalisation.
What else the public tests measure
- tasksource held-out (2,206 questions, 350 sources): new examples of tasks seen in training. Every source of this test is also in the Stage 2 mix (from its train split); no example group is shared. Accuracy 49.5 % [47.4–51.6] (Stage 1: 38.5 %; paired p = 1e-14). It is not an evaluation on tasks absent from training; for that, see BTZSC-22 below.
- No typed-decisions test case, state, normalised state text or state structure appears in any training file.
External figures, kept apart (no leaderboard ranking)
General / zero-shot models, as published (measured by others, whole case per request):
| model | accuracy | Brier | source |
|---|---|---|---|
| TypeSafe Jev 1.13 | 74.0 [72.1–75.9] | 0.148 | DecisionEval, independent run, test frozen 2026-09-20 |
Ines-1 is not ranked against Jev. Its 76.4 % is a single-question specialist result; Jev's 74.0 % is zero-shot on
whole cases. Specialists trained on typed-decisions train have a separate table in the dataset card (mostly
self-reported, 0.483 to 0.801); the card also notes that scores above its teacher self-agreement (0.735) indicate
learning the teacher's quirks.
External zero-shot classification: BTZSC-22
We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.
Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1, 0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for Julia-1 (+0.051, 95% CI −0.026 to +0.126) or laya-multilingual (+0.056, 95% CI −0.040 to +0.163). On the 15 datasets absent from Ines-1's decision fine-tuning mixture, none of the pairwise differences was resolved.
For datasets with at most 20 labels, native ECE was 0.127 for Ines-1, 0.298 for Julia-1, 0.093 for laya-multilingual, and 0.159 for GLiNER2.5.
We also tested alternative decompositions of multiclass decisions. A common one-vs-rest protocol did not improve Ines-1 (mean macro-F1 0.457 overall, and 0.489 → 0.283 on the 3–20-label subset). In contrast, a descriptive pairwise-knockout ablation improved Ines-1 on 6 of 9 multiclass datasets (accuracy). This suggests that Ines-1's current interface benefits more from comparing a small set of competing alternatives than from independently scoring each label as a yes/no decision. The knockout result is order-dependent and is reported as an ablation rather than a general classification protocol.
Full protocol, manifests, prediction hashes, per-dataset results, calibration details, and reproducibility notes are
available in eval/btzsc22/.
Private domain benchmark
On an internal held-out domain benchmark (9 Spanish subtasks, 1,026 questions; not released), generic typed-decision training does not transfer uniformly zero-shot. Macro balanced accuracy: Pretrain-V3 0.280, Stage 1 0.447, Stage 2 from pretraining 0.543, Ines-1 zero-shot 0.558, an untuned 322M multilingual encoder 0.553, an untuned 144M encoder 0.485. On one imbalanced yes/no subtask no untuned checkpoint separates the classes (AUROC 0.48–0.55).
Fine-tuning this exact released checkpoint on that domain (recipe of the earlier domain checkpoints, only the parent changed; epoch 3 chosen on validation) raised macro balanced accuracy 0.558 → 0.790 (micro accuracy 0.788, macro-F1 0.780; Stage 1 + the same domain fine-tuning: 0.784); on the yes/no subtask AUROC 0.541 → 0.889; organisations never seen in training scored 0.860 accuracy (0.678 balanced) against 0.841 (0.665) for seen ones. Cost: 21,423 domain cases, 23,608 questions per epoch, 28.5M prompt tokens, 2.25 H200 GPU-hours. Specialisation cost 6–9 points of typed-decisions accuracy: EN 0.764 → 0.698, ES 0.754 → 0.689, independent ES 0.763 → 0.675 (tasksource 0.495 → 0.480). The domain-adapted model is private and is not included in this repository.
Ablations (public tests, same questions, exact McNemar)
| comparison | typed EN | typed ES (ours) | typed ES (indep.) | tasksource |
|---|---|---|---|---|
| Stage 1 → release | 75.5 → 76.4 (p = 0.24) | 74.7 → 75.4 (p = 0.32) | 74.6 → 76.3 (p = 0.023) | 38.5 → 49.5 (p = 1e-14) |
| release, Engram removed at inference | 76.4 → 74.7 (p = 0.018) | 75.4 → 72.8 (p = 0.001) | 76.3 → 73.5 (p = 3e-4) | 49.5 → 47.0 (p = 0.007) |
| Stage 2 from pretraining (instead of from Stage 1) | 72.6 | 70.4 | 71.2 | 50.7 |
"Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.
Speed
eval/speed.md has the full protocol and both runs. One H200, bf16, HTTP closed loop, same requests and load generator for every model. Two regimes, kept apart:
| single request p50 / p95 | max throughput (questions/s) | |
|---|---|---|
| short request (2 questions, ~60 prompt tokens each) — Ines-1, 1 process | 9.6 / 9.7 ms | ~1,900 |
| same, 144M / 322M encoders, 4 processes each | 12.7 / 21–26 ms | ~1,700–1,850 |
| typed-decisions case (5 questions, ~315 tokens each) — Ines-1 | 29.6 / 38.3 ms | ~540 |
| same, 144M / 322M encoders | 15–16 / 87–88 ms | ~760–1,110 |
On ~300-token questions the encoders are 1.4–2.1× faster.
Training
| stage | parent weights (sha256) | data | epochs | optimizer steps | output weights (fp32, sha256) |
|---|---|---|---|---|---|
| Pretrain-V3 | — | 9.96B tokens processed from the ~10B-token corpus below | 1 pass | 152,588 × 65,536 tokens | dd3a794d…57f9 |
| Decision-FT Stage 1 | dd3a794d… |
typed-decisions train EN + our ES translation + 9.5k jev-decisions: 20,233 questions | 8 scheduled, 7 effective | 8,852 of 10,117 | d5c76527…8fbf |
| Decision-Mix Stage 2 (Ines-1) | d5c76527… |
balanced mix, 86,436 questions | 2 | 10,805 | 5657d3ee…e8fb |
The released model.safetensors is the Stage 2 fp32 checkpoint cast to bf16.
Recipe for both decision stages: full fine-tuning, AdamW lr 1e-5 (β 0.9/0.95, no weight decay, clip 1.0), accumulation 16 (one question per forward), linear warm-up over 5 % of the scheduled steps then linear decay to 0, cross-entropy of the letter distribution against the teacher distribution when present (one-hot otherwise) + 0.01 × MoE auxiliary loss, choice options shuffled per epoch, epoch chosen on the mean of one validation file per data family, seed 7. The typed-decisions training questions are seen in both stages (7 + 2 epochs).
Code provenance, stated as it happened: Stage 2 ran from commit 0b5e63e with uncommitted files (-dirty: the
mix builder prepare_mix.py was not yet committed). It was committed afterwards (039b99c) and rebuilding the mix from
that clean commit reproduces all 16 data files byte for byte. Bundle code comes only from commits (see config.json).
Training data
Full provenance, identifiers and licence notes: THIRD_PARTY_DATA_NOTICE.md; per-source licences of the tasksource questions used: TASKSOURCE_LICENSE_AUDIT.md; how our Spanish translation was made: training/TRANSLATION_PROVENANCE.md. No training data is redistributed in this repository.
Pretraining (Pretrain-V3, ~10B tokens, realised mix)
| domain | share | sources |
|---|---|---|
| general English | 46.5 % | FineWeb-Edu, FineWeb |
| Spanish | 20.0 % | FineWeb-2 (spa_Latn) |
| reference | 15.0 % | Common Pile: Wikimedia, DOAB, peS2o, Python Enhancement Proposals |
| math | 9.0 % | OpenWebMath, Common Pile StackExchange (math sites) |
| code | 8.5 % | Common Pile Stack-Edu |
| structured | 1.0 % | 75M tokens of JSON/YAML/XML mined from Stack-Edu + 25M synthetic |
Approximately 25M pretraining tokens (~0.25% of the corpus) were synthetic structured documents (JSON-schema style documents generated for this corpus; not redistributed). Their contribution was not isolated.
Decision fine-tuning (Stage 1 and Stage 2)
Stage 1: typed-decisions train (English) + our Spanish translation of it + 9.5k jev-decisions questions. Stage 2 mix (no family above ~29 %):
| family | questions | source |
|---|---|---|
| typed-decisions EN | 5,420 | LocalLLaMA/typed-decisions @ c76749ec (data files identical to the 2026-09-18 revision) |
| typed-decisions ES (ours) | 5,335 | our machine translation of the above (text values only; keys and structure checked; provenance) |
| typed-decisions ES (independent) | 5,415 | telepatia-ai/typed-decisions-pt-es @ 8c85a4d2, Spanish only; 117 cases matching our validation split removed |
| tool selection | 17,898 | samatv256/jev-decisions-v1 @ c12aadf1, general-clean-50k (CC-BY-4.0) |
| tasksource | 24,959 | tasksource/tasksource-jev-typed-decisions @ 8173a06c: rows whose metadata says license_use=commercial, ≤ 80 per source, 366 sources, none derived from typed- or jev-decisions |
| synthetic, teacher soft labels | 18,001 | n4ze3m/typed-decisions-synth @ 5ece89a2 |
| synthetic | 9,408 | helmo/synthetic-typed-decisions @ 1827dc0d |
Limitations
- One-question interface. With the whole case in one prompt, accuracy falls to 37–42 % (below the dataset's Prior); API batching does not change that the model sees one question at a time.
- Specialist exposure to the benchmark. The typed-decisions results are in-distribution for a model trained on its train split; they are not comparable with zero-shot results.
- Limited evaluation on tasks absent from training. tasksource measures new examples of tasks seen in training; BTZSC-22 includes 15 datasets absent from the decision fine-tuning mixture (100 examples each; pretraining exposure not excluded), where no difference against the three classifiers is resolved.
- Uneven zero-shot domain transfer (0.558 macro balanced accuracy on a private domain), including a class-prior shift on yes/no questions over long texts after Stage 2.
- Specialisation forgetting. Domain fine-tuning cost 6–9 points of typed-decisions accuracy.
- No retrieval. 2,048-token context; it fails a passkey retrieval test even at 1k tokens: it decides, it does not look things up.
- Engram reliance was measured only by removal at inference; no model was trained without Engram.
- Pretrain-V2 → V3 is not a pure token-scaling ablation: corpus, horizon and schedule changed together (and V3 is the first corpus with synthetic documents).
- Soft vs hard labels not ablated; the contribution of teacher distributions is not isolated.
- Synthetic contribution not isolated, in pretraining or in fine-tuning.
- our-ECE is not comparable with published ECE values; probabilities are calibrated by training, not certified.
- Weaker on
choicethan onscoreandnoul. Spanish results are on machine translations (two independent translations agree within 1 point for every model). Not evaluated for safety-critical decisions.
Reproducibility
- Environment: ENVIRONMENT.md (GPU, driver, Python, PyTorch, settings),
requirements.txt(minimum),requirements.lock(exact versions of the evaluation environment),Dockerfile(that environment as an image; built and smoke-tested, see ENVIRONMENT.md). File hashes:SHA256SUMS;model.safetensorssha2567fa38ab3adf37530c5b150b0527d8cc9c24350b402cd2d8bf727e736b78ac555. - Canonical inference (every number in this card): bf16 weights under
torch.autocast("cuda", dtype=torch.bfloat16), one prompt per question,prefill(num_logits=1)— the default path ofDecider.decide. - Canonical evaluator:
training/code/eval_items.py(per-item outputs) andtraining/code/metrics.py; their outputs for this model are ineval/items/andeval/report.json. - HTTP / batched path (
scripts/serve_jev.py,Decider.decide(..., batched=True)): right-padded rows in one prefill, with CUDA graphs on the server. It is not bit-identical to the canonical path and may change a few answers because of batching and bf16 numerics (8 of 400 on a probe). Through HTTP the English test scored 1527 to 1531 / 2000 in three runs of the same server (1975–1976 of 2000 answers equal to the canonical ones; concurrent batching changes from run to run), vs 1528 / 2000 canonical. - PyTorch 2.13 and 2.14 gave bit-identical outputs with the same code; autocast on/off and the attention backend each changed 8 of 400 answers (ENVIRONMENT.md).
Future work (hypotheses, not results)
Branched / tree attention or a shared-state prefill so that several questions are answered independently in one forward; training explicitly on multi-question cases; and measuring whether that restores full-case behaviour without sacrificing single-question performance.
Files
model.safetensors (bf16), config.json (configuration + provenance), tokenizer/, mini_v41/ (modelling code),
mini_v41_jev/ (reading decisions, Decider), scripts/ (HTTP server, load test), examples/, training/ (recipe,
manifests, data audit, scripts), eval/ (reports, per-item outputs, speed protocol, btzsc22/ external classification evaluation), ENVIRONMENT.md,
requirements.txt, requirements.lock, Dockerfile, CHANGELOG_CLAIMS.md,
THIRD_PARTY_DATA_NOTICE.md, LICENSING_NOTES.md,
TASKSOURCE_LICENSE_AUDIT.md + tasksource_license_audit.csv,
paper/ (technical report: PDF + LaTeX sources), LICENSE (Apache-2.0), NOTICE, LICENSE-CODE (MIT), SHA256SUMS.
Paper
The technical report is available in this repository:
Ines-1: Training a 1.6B Mixture-of-Experts Model for Typed Decisions from Scratch
An arXiv submission is planned; the citation will be updated once an arXiv identifier is available. The LaTeX sources
(paper/main.tex, paper/references.bib, paper/figures/) build the same PDF with pdflatex + bibtex.
Citation
Preprint; the identifier will be added when it exists.
@misc{endikavi2026ines1,
title = {Ines-1: Training a 1.6B Mixture-of-Experts Model for Typed Decisions from Scratch},
author = {Endikavi},
year = {2026},
note = {Preprint}
}
- Downloads last month
- -