Ines-1

Licences — what covers what.

  • Apache-2.0 (LICENSE, NOTICE): the model weights (model.safetensors), config.json and tokenizer/; the documentation (README.md and the other .md files, paper/, tasksource_license_audit.csv); the training records (training/*.json); the model's own evaluation outputs in eval/; SHA256SUMS.
  • MIT (LICENSE-CODE): the code — mini_v41/, mini_v41_jev/, scripts/, examples/, training/code/, eval/btzsc22/code/ — and the environment and build files: Dockerfile, .dockerignore, requirements.txt, requirements.lock, .gitattributes, .gitignore.
  • Third-party terms, not relicensed: the gold labels and teacher distributions of the public test sets that eval/items/, eval/reference_rows.json and eval/btzsc22/ (BTZSC label texts and targets) include keep their datasets' licences. The training data is not redistributed and remains under its own terms, including attribution, share-alike and copyleft ones: see THIRD_PARTY_DATA_NOTICE.md and LICENSING_NOTES.md.

Ines-1 is a typed-decision model pretrained from scratch: a 1.59B-parameter mixture-of-experts causal encoder–decoder (405M parameters active per token) with hashed n-gram memory (Engram), pretrained on about 10B public tokens on one H200, then fine-tuned to answer one typed question at a time — choice, score or noul (yes/no) — with a probability for every option, in English and Spanish. The request shape follows the typed-decision POST /v1/systemone interface popularised by TypeSafe's Jev. Ines-1 does not reproduce Jev's architecture or its full-case (multi-question) protocol.

Development codename: mini-v41 / mini-v41-Decisions. The Python packages keep that name (mini_v41, mini_v41_jev), and so do the provenance records in config.json and training/.

Supported decision interface

  • One typed question per model prompt. The answer is the softmax over the option letters (at most 26) at the first reply position.
  • Several questions of a case can be API-batched: the server (and Decider.decide) splits a case into one prompt per question and scores them together. This is batching, not multi-question modelling.
  • The model was not trained to condition jointly on several typed questions in one prompt, and it does not do so well: with the whole case in the prompt its accuracy falls from 76.4 % to 37–42 % (Full-context stress tests).

Intended use

  • Typed single-question decisions over a piece of state (text or JSON): classification, routing and scoring over a closed, caller-defined option set, with a probability per option.
  • A starting checkpoint for further domain specialisation (fine-tuning on your own decisions).

Out of scope, or not supported by our evidence: use as a chat or general assistant; retrieval or fact lookup in long documents; native multi-question (Jev full-case) compatibility; strong zero-shot performance on a new domain; any safety-critical decision.

Parameters 1,590,066,240 total; 405,052,224 active per token; weights in bf16
Layers 8 encoder + 8 decoder in one causal stack; hidden 1024; 16 heads / 4 KV heads, head dim 64
Decoder sliding window of 128 + one global K/V projected from the encoder output; receptive field 1,017 positions
MoE every layer: 16 routed experts (FFN 1536), top-2, + 1 shared expert
Engram hashed 2/3/4-gram embeddings at encoder layers 1 and 5 (2 × 10⁶ rows, 128M parameters)
Context 2,048 tokens; longer cases are trimmed in the middle
Tokenizer byte-level BPE, 65,536 tokens (identity sha256 859e5c63…74469, checked at load)

Quickstart

pip install -r requirements.txt   # requirements.lock pins the canonical environment (ENVIRONMENT.md)
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("Endikavi/Ines-1")   # the whole repository: weights, tokenizer and reader code
sys.path.insert(0, path)
from mini_v41_jev import Decider

decider = Decider(path)                       # CUDA if available (bf16 + autocast), otherwise CPU (fp32)
question = {"type": "choice",
            "instructions": "Which department should handle this ticket?",
            "criteria": {"billing": "payments and invoices", "tech": "outages and bugs", "legal": "disputes"}}
state = "The invoice was charged twice and the customer wants the money back today."

answer = decider.decide(state, {"dept": question}, lang="en")["dept"]
print(answer["choice"])          # argmax option
print(answer["probabilities"])   # probability of every option

score questions return probabilities over levels "0".."n-1"; noul questions return {"noul": P(true)}. lang="es" uses Spanish prompts. examples/quickstart.py runs the same from a local copy.

HTTP server with the same request shape (POST /v1/systemone, GET /v1/models, GET /health; CUDA and aiohttp), batching concurrent questions into one prefill: python scripts/serve_jev.py --port 30040.

Lineage

Pretrain-V3 (9.96B tokens) → Decision-FT Stage 1 (typed-decisions + jev-decisions; 8 epochs scheduled, 7 effective: the epoch chosen on validation is the checkpoint kept) → Decision-Mix Stage 2 (balanced public mix, 2 epochs) → Ines-1. Each stage's parent-weights hash matches the previous stage's output (table under Training).

Results

All numbers measured by us on the released weights in the canonical environment (ENVIRONMENT.md: bf16 weights under bf16 autocast). Per-item outputs in eval/items/, reports in eval/. Accuracy = the highest-probability option equals the gold label; intervals are bootstrap over cases (the five questions of a case are not independent).

typed-decisions — single-question specialist protocol

Single-question specialist protocol, after training on the benchmark's training split: one prompt per question; the model was trained on typed-decisions train (same four workflows, with the dataset's teacher distributions). This is not a zero-shot result.

test (400 cases) correct / n accuracy 95 % CI choice score noul Brier KL(gold‖p) log loss
English 1528 / 2000 76.4 74.2–78.5 71.5 74.5 83.8 0.065 0.121 0.635
Spanish (our machine translation) 1496 / 1985 75.4 73.2–77.4 69.9 73.9 82.7 0.065 0.119 0.639
Spanish (independent translation) 1526 / 2000 76.3 74.2–78.5 73.0 73.8 83.0 0.064 0.119 0.639

Brier and KL are against the teacher's gold distribution and follow the dataset card's definitions (they reproduce its Uniform reference row exactly: KL 0.444, Brier 0.238). Our ECE (10 equal-width bins of top-1 confidence; 0.132 on English) is reported as our-ECE and is not comparable with published ECE values, whose code is not public. By teacher agreement (label_agreement.argmax_agree): 86.4 % where the teacher's samples agree (n = 1,188), 61.7 % where they do not (n = 812).

typed-decisions — full-context stress tests

The dataset's public runs send the whole case (state + its five questions) in one request, and its card notes that request shape can change results. Ines-1 has no readout that answers several questions in one forward, so we ran two stress tests instead (English; Spanish behaves the same):

readout what the model sees and how answers are read correct / n accuracy 95 % CI Brier
single question (above) the state and one question 1528 / 2000 76.4 74.2–78.5 0.065
context the state and all five questions, ending "Answer question k"; one forward per question 831 / 2000 41.5 39.5–43.6 0.266
joint one prompt with all five questions; answers read in sequence as "k: letter", each conditioned on the model's previous answers 742 / 2000 37.1 34.9–39.3 0.246
dataset reference: Prior train label frequencies, ignoring the input 47.0 0.189

Neither stress test reproduces the public protocol or a Jev-style single forward over five independent decisions; neither is a DecisionEval score. Both show strong sensitivity to request shape. Decision-FT Stage 1 drops the same way (75.5 → 42.8 %, context), so this is not caused by Stage 2. Task competence does not imply interface generalisation.

What else the public tests measure

  • tasksource held-out (2,206 questions, 350 sources): new examples of tasks seen in training. Every source of this test is also in the Stage 2 mix (from its train split); no example group is shared. Accuracy 49.5 % [47.4–51.6] (Stage 1: 38.5 %; paired p = 1e-14). It is not an evaluation on tasks absent from training; for that, see BTZSC-22 below.
  • No typed-decisions test case, state, normalised state text or state structure appears in any training file.

External figures, kept apart (no leaderboard ranking)

General / zero-shot models, as published (measured by others, whole case per request):

model accuracy Brier source
TypeSafe Jev 1.13 74.0 [72.1–75.9] 0.148 DecisionEval, independent run, test frozen 2026-09-20

Ines-1 is not ranked against Jev. Its 76.4 % is a single-question specialist result; Jev's 74.0 % is zero-shot on whole cases. Specialists trained on typed-decisions train have a separate table in the dataset card (mostly self-reported, 0.483 to 0.801); the card also notes that scores above its teacher self-agreement (0.735) indicate learning the teacher's quirks.

External zero-shot classification: BTZSC-22

We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.

Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1, 0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for Julia-1 (+0.051, 95% CI −0.026 to +0.126) or laya-multilingual (+0.056, 95% CI −0.040 to +0.163). On the 15 datasets absent from Ines-1's decision fine-tuning mixture, none of the pairwise differences was resolved.

For datasets with at most 20 labels, native ECE was 0.127 for Ines-1, 0.298 for Julia-1, 0.093 for laya-multilingual, and 0.159 for GLiNER2.5.

We also tested alternative decompositions of multiclass decisions. A common one-vs-rest protocol did not improve Ines-1 (mean macro-F1 0.457 overall, and 0.489 → 0.283 on the 3–20-label subset). In contrast, a descriptive pairwise-knockout ablation improved Ines-1 on 6 of 9 multiclass datasets (accuracy). This suggests that Ines-1's current interface benefits more from comparing a small set of competing alternatives than from independently scoring each label as a yes/no decision. The knockout result is order-dependent and is reported as an ablation rather than a general classification protocol.

Full protocol, manifests, prediction hashes, per-dataset results, calibration details, and reproducibility notes are available in eval/btzsc22/.

Private domain benchmark

On an internal held-out domain benchmark (9 Spanish subtasks, 1,026 questions; not released), generic typed-decision training does not transfer uniformly zero-shot. Macro balanced accuracy: Pretrain-V3 0.280, Stage 1 0.447, Stage 2 from pretraining 0.543, Ines-1 zero-shot 0.558, an untuned 322M multilingual encoder 0.553, an untuned 144M encoder 0.485. On one imbalanced yes/no subtask no untuned checkpoint separates the classes (AUROC 0.48–0.55).

Fine-tuning this exact released checkpoint on that domain (recipe of the earlier domain checkpoints, only the parent changed; epoch 3 chosen on validation) raised macro balanced accuracy 0.558 → 0.790 (micro accuracy 0.788, macro-F1 0.780; Stage 1 + the same domain fine-tuning: 0.784); on the yes/no subtask AUROC 0.541 → 0.889; organisations never seen in training scored 0.860 accuracy (0.678 balanced) against 0.841 (0.665) for seen ones. Cost: 21,423 domain cases, 23,608 questions per epoch, 28.5M prompt tokens, 2.25 H200 GPU-hours. Specialisation cost 6–9 points of typed-decisions accuracy: EN 0.764 → 0.698, ES 0.754 → 0.689, independent ES 0.763 → 0.675 (tasksource 0.495 → 0.480). The domain-adapted model is private and is not included in this repository.

Ablations (public tests, same questions, exact McNemar)

comparison typed EN typed ES (ours) typed ES (indep.) tasksource
Stage 1 → release 75.5 → 76.4 (p = 0.24) 74.7 → 75.4 (p = 0.32) 74.6 → 76.3 (p = 0.023) 38.5 → 49.5 (p = 1e-14)
release, Engram removed at inference 76.4 → 74.7 (p = 0.018) 75.4 → 72.8 (p = 0.001) 76.3 → 73.5 (p = 3e-4) 49.5 → 47.0 (p = 0.007)
Stage 2 from pretraining (instead of from Stage 1) 72.6 70.4 71.2 50.7

"Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.

Speed

eval/speed.md has the full protocol and both runs. One H200, bf16, HTTP closed loop, same requests and load generator for every model. Two regimes, kept apart:

single request p50 / p95 max throughput (questions/s)
short request (2 questions, ~60 prompt tokens each) — Ines-1, 1 process 9.6 / 9.7 ms ~1,900
same, 144M / 322M encoders, 4 processes each 12.7 / 21–26 ms ~1,700–1,850
typed-decisions case (5 questions, ~315 tokens each) — Ines-1 29.6 / 38.3 ms ~540
same, 144M / 322M encoders 15–16 / 87–88 ms ~760–1,110

On ~300-token questions the encoders are 1.4–2.1× faster.

Training

stage parent weights (sha256) data epochs optimizer steps output weights (fp32, sha256)
Pretrain-V3 — 9.96B tokens processed from the ~10B-token corpus below 1 pass 152,588 × 65,536 tokens dd3a794d…57f9
Decision-FT Stage 1 dd3a794d… typed-decisions train EN + our ES translation + 9.5k jev-decisions: 20,233 questions 8 scheduled, 7 effective 8,852 of 10,117 d5c76527…8fbf
Decision-Mix Stage 2 (Ines-1) d5c76527… balanced mix, 86,436 questions 2 10,805 5657d3ee…e8fb

The released model.safetensors is the Stage 2 fp32 checkpoint cast to bf16.

Recipe for both decision stages: full fine-tuning, AdamW lr 1e-5 (β 0.9/0.95, no weight decay, clip 1.0), accumulation 16 (one question per forward), linear warm-up over 5 % of the scheduled steps then linear decay to 0, cross-entropy of the letter distribution against the teacher distribution when present (one-hot otherwise) + 0.01 × MoE auxiliary loss, choice options shuffled per epoch, epoch chosen on the mean of one validation file per data family, seed 7. The typed-decisions training questions are seen in both stages (7 + 2 epochs).

Code provenance, stated as it happened: Stage 2 ran from commit 0b5e63e with uncommitted files (-dirty: the mix builder prepare_mix.py was not yet committed). It was committed afterwards (039b99c) and rebuilding the mix from that clean commit reproduces all 16 data files byte for byte. Bundle code comes only from commits (see config.json).

Training data

Full provenance, identifiers and licence notes: THIRD_PARTY_DATA_NOTICE.md; per-source licences of the tasksource questions used: TASKSOURCE_LICENSE_AUDIT.md; how our Spanish translation was made: training/TRANSLATION_PROVENANCE.md. No training data is redistributed in this repository.

Pretraining (Pretrain-V3, ~10B tokens, realised mix)

domain share sources
general English 46.5 % FineWeb-Edu, FineWeb
Spanish 20.0 % FineWeb-2 (spa_Latn)
reference 15.0 % Common Pile: Wikimedia, DOAB, peS2o, Python Enhancement Proposals
math 9.0 % OpenWebMath, Common Pile StackExchange (math sites)
code 8.5 % Common Pile Stack-Edu
structured 1.0 % 75M tokens of JSON/YAML/XML mined from Stack-Edu + 25M synthetic

Approximately 25M pretraining tokens (~0.25% of the corpus) were synthetic structured documents (JSON-schema style documents generated for this corpus; not redistributed). Their contribution was not isolated.

Decision fine-tuning (Stage 1 and Stage 2)

Stage 1: typed-decisions train (English) + our Spanish translation of it + 9.5k jev-decisions questions. Stage 2 mix (no family above ~29 %):

family questions source
typed-decisions EN 5,420 LocalLLaMA/typed-decisions @ c76749ec (data files identical to the 2026-09-18 revision)
typed-decisions ES (ours) 5,335 our machine translation of the above (text values only; keys and structure checked; provenance)
typed-decisions ES (independent) 5,415 telepatia-ai/typed-decisions-pt-es @ 8c85a4d2, Spanish only; 117 cases matching our validation split removed
tool selection 17,898 samatv256/jev-decisions-v1 @ c12aadf1, general-clean-50k (CC-BY-4.0)
tasksource 24,959 tasksource/tasksource-jev-typed-decisions @ 8173a06c: rows whose metadata says license_use=commercial, ≤ 80 per source, 366 sources, none derived from typed- or jev-decisions
synthetic, teacher soft labels 18,001 n4ze3m/typed-decisions-synth @ 5ece89a2
synthetic 9,408 helmo/synthetic-typed-decisions @ 1827dc0d

Limitations

  • One-question interface. With the whole case in one prompt, accuracy falls to 37–42 % (below the dataset's Prior); API batching does not change that the model sees one question at a time.
  • Specialist exposure to the benchmark. The typed-decisions results are in-distribution for a model trained on its train split; they are not comparable with zero-shot results.
  • Limited evaluation on tasks absent from training. tasksource measures new examples of tasks seen in training; BTZSC-22 includes 15 datasets absent from the decision fine-tuning mixture (100 examples each; pretraining exposure not excluded), where no difference against the three classifiers is resolved.
  • Uneven zero-shot domain transfer (0.558 macro balanced accuracy on a private domain), including a class-prior shift on yes/no questions over long texts after Stage 2.
  • Specialisation forgetting. Domain fine-tuning cost 6–9 points of typed-decisions accuracy.
  • No retrieval. 2,048-token context; it fails a passkey retrieval test even at 1k tokens: it decides, it does not look things up.
  • Engram reliance was measured only by removal at inference; no model was trained without Engram.
  • Pretrain-V2 → V3 is not a pure token-scaling ablation: corpus, horizon and schedule changed together (and V3 is the first corpus with synthetic documents).
  • Soft vs hard labels not ablated; the contribution of teacher distributions is not isolated.
  • Synthetic contribution not isolated, in pretraining or in fine-tuning.
  • our-ECE is not comparable with published ECE values; probabilities are calibrated by training, not certified.
  • Weaker on choice than on score and noul. Spanish results are on machine translations (two independent translations agree within 1 point for every model). Not evaluated for safety-critical decisions.

Reproducibility

  • Environment: ENVIRONMENT.md (GPU, driver, Python, PyTorch, settings), requirements.txt (minimum), requirements.lock (exact versions of the evaluation environment), Dockerfile (that environment as an image; built and smoke-tested, see ENVIRONMENT.md). File hashes: SHA256SUMS; model.safetensors sha256 7fa38ab3adf37530c5b150b0527d8cc9c24350b402cd2d8bf727e736b78ac555.
  • Canonical inference (every number in this card): bf16 weights under torch.autocast("cuda", dtype=torch.bfloat16), one prompt per question, prefill(num_logits=1) — the default path of Decider.decide.
  • Canonical evaluator: training/code/eval_items.py (per-item outputs) and training/code/metrics.py; their outputs for this model are in eval/items/ and eval/report.json.
  • HTTP / batched path (scripts/serve_jev.py, Decider.decide(..., batched=True)): right-padded rows in one prefill, with CUDA graphs on the server. It is not bit-identical to the canonical path and may change a few answers because of batching and bf16 numerics (8 of 400 on a probe). Through HTTP the English test scored 1527 to 1531 / 2000 in three runs of the same server (1975–1976 of 2000 answers equal to the canonical ones; concurrent batching changes from run to run), vs 1528 / 2000 canonical.
  • PyTorch 2.13 and 2.14 gave bit-identical outputs with the same code; autocast on/off and the attention backend each changed 8 of 400 answers (ENVIRONMENT.md).

Future work (hypotheses, not results)

Branched / tree attention or a shared-state prefill so that several questions are answered independently in one forward; training explicitly on multi-question cases; and measuring whether that restores full-case behaviour without sacrificing single-question performance.

Files

model.safetensors (bf16), config.json (configuration + provenance), tokenizer/, mini_v41/ (modelling code), mini_v41_jev/ (reading decisions, Decider), scripts/ (HTTP server, load test), examples/, training/ (recipe, manifests, data audit, scripts), eval/ (reports, per-item outputs, speed protocol, btzsc22/ external classification evaluation), ENVIRONMENT.md, requirements.txt, requirements.lock, Dockerfile, CHANGELOG_CLAIMS.md, THIRD_PARTY_DATA_NOTICE.md, LICENSING_NOTES.md, TASKSOURCE_LICENSE_AUDIT.md + tasksource_license_audit.csv, paper/ (technical report: PDF + LaTeX sources), LICENSE (Apache-2.0), NOTICE, LICENSE-CODE (MIT), SHA256SUMS.

Paper

The technical report is available in this repository:

Ines-1: Training a 1.6B Mixture-of-Experts Model for Typed Decisions from Scratch

An arXiv submission is planned; the citation will be updated once an arXiv identifier is available. The LaTeX sources (paper/main.tex, paper/references.bib, paper/figures/) build the same PDF with pdflatex + bibtex.

Citation

Preprint; the identifier will be added when it exists.

@misc{endikavi2026ines1,
  title  = {Ines-1: Training a 1.6B Mixture-of-Experts Model for Typed Decisions from Scratch},
  author = {Endikavi},
  year   = {2026},
  note   = {Preprint}
}
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Endikavi/Ines-1