matilda-jev-v1 / README.md
yue-maincode's picture
Update README.md
629e8fe verified
|
Raw History Blame Contribute Delete
12.3 kB
metadata
pretty_name: Matilda Jev
license: apache-2.0
license_name: matilda-jev
language:
  - en
library_name: transformers
tags:
  - decision-model
  - typed-decisions
  - jev
  - decision-index
  - maincode
  - australia

Matilda Jev

TL;DR. Matilda Jev is Australia's first Jev-style decision model, conditioned for Australian compliance and political even-handedness, then post-trained with our own specialised recipe. Scored with the official Jev Decision Index kit, it reaches a score of 59.26, beating Jev which scored 57.91.

Much of the work people hand to language models is not open-ended writing. It is decisions: which tool to call, which label applies, which document answers the question, whether a claim is supported. Jev (Typesafe.ai) showed that these are best served by a model that answers them directly, as probabilities, rather than by parsing generated text. Matilda Jev is Maincode's model of that kind, built so that it behaves like a model deployed in Australia without giving up capability.

Developer Maincode, Australia
Head Same as Jev, 255-way answer-code readout over the backbone's final hidden state (1.3M parameters)
Output A probability for every option of a choice, noul (yes/no) or score question, temperature-calibrated
Precision / size bf16, ~49 GiB of weights
Context Up to 131k tokens per question on GPUs with ≥ 200 GB, 32k on 96 GB cards (the server sizes this automatically)
Inputs Text or JSON state; optional images (up to 4)
Languages English

How it works

Each question is rendered into one prompt (state, question, lettered options, "return only the letter code") with thinking disabled. The model reads it once. A readout scores the answer codes A, B, … (up to 255 options, each a single token), and a stored temperature turns those scores into calibrated probabilities. There is no sampling and no answer parsing, so the output is deterministic and every option always gets a probability.

  • choice: pick among named options (1–255). Returns the distribution and the most likely option.
  • noul: a yes/no question. Returns P(yes).
  • score: an ordered rubric (2–10 levels). Returns the distribution and its expected level.

A request can carry several questions over one shared state; each is one forward pass, batched together.

Results

Decision Index 0.2.1, scored with the official kit and scorer over the full suite (150,317 requests; every request answered, none truncated). Chance-corrected: 0 = random guessing, 100 = perfect. Board entries as published on the leaderboard on 2026-09-28.

Model Params Decision Index Raw Breadth
Matilda Jev 26.1B 59.26 68.89 58.06
Jev (Typesafe.ai, hosted) – 57.91 68.09 57.08
Surogate Rune 26B-A4B v3 25.8B 57.44 67.30 56.47
Decider chat 32.7B 57.33 67.22 56.20
AutoJev-27B 27.8B 56.40 66.89 54.93
simple-jev 27.8B 55.74 66.26 53.91
Jebadiah 27B 27.8B 54.67 65.62 53.18
Eikos-27B 27.8B 53.13 63.64 52.02

By area (chance-corrected skill × 100):

Model Knowledge & Reasoning Language Retrieval & Classification Tools & Automation Arts & Human Taste
Matilda Jev 43.52 66.36 61.34 77.72 43.65
Jev 51.40 62.02 55.42 75.09 37.66
Surogate Rune 26B-A4B v3 43.36 63.06 63.52 71.22 41.90
AutoJev-27B 40.88 63.46 54.86 79.35 39.39

Latency. Median 56.6 ms per request across the suite on 1× AMD Instinct MI355X, bf16, one request at a time. The board re-measures latency on its own RTX PRO 6000.

Contamination. Every training row whose text copies a Decision Index test item was removed before training (exact-copy check against all 155,390 suite requests: 0 remaining). We are submitting the model to the board so the score can be verified with its own hardware, contamination and option-shuffle checks.

Hardware. Trained on one node of 8× AMD Instinct MI355X (288 GB each) in Maincode's MC-2 cluster in Australia, with full-weight bf16 training.

Run

Python 3.12+, uv, and a GPU with room for ~49 GiB of bf16 weights plus runtime overhead (80 GB or more recommended; tested on AMD MI355X, targeted at NVIDIA RTX PRO 6000).

git clone https://github.com/Maincode/matilda-jev.git
cd matilda-jev
uv sync --frozen --python 3.12 --extra cuda      # NVIDIA (driver >= 580); use --extra rocm on AMD
uv run --no-sync hf download Maincode/matilda-jev --local-dir checkpoints/matilda-jev
MJ_CHECKPOINT=checkpoints/matilda-jev uv run --no-sync matilda-jev-serve

The API is at http://localhost:8000 (interactive docs at /docs). POST /v1/systemone supports choice, noul, score, and optional base64 images. Set MJ_API_KEY (or MJ_API_KEY_FILE) to enable authentication and MJ_HOST=0.0.0.0 to listen beyond localhost. While the weights are private, authenticate with uv run --no-sync hf auth login before downloading.

Alternatives: scripts/serve.sh [--background] picks the CUDA or ROCm stack for you and reads serve.env; docker compose up --build runs it on NVIDIA with ./checkpoints mounted read-only.

Request

curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "matilda-jev-latest",
  "state": {"ticket": "My card was charged twice for the same order."},
  "questions": {
    "intent": {"type": "choice", "instructions": "Main issue?",
               "criteria": {"duplicate_charge": "Charged twice for one purchase", "card_delivery": null, "other": null}},
    "urgent": {"type": "noul", "instructions": "Does this need a same-day response?"},
    "anger":  {"type": "score", "instructions": "How frustrated is the customer?",
               "criteria": ["calm", "annoyed", "frustrated", "angry"]}
  }
}'
{"model": "matilda-jev",
 "answers": {
   "intent": {"type": "choice", "choice": "duplicate_charge", "probabilities": {"duplicate_charge": 0.97, "card_delivery": 0.01, "other": 0.02}, "confidence": 0.97},
   "urgent": {"type": "noul", "noul": 0.81},
   "anger":  {"type": "score", "score": 1.6, "probabilities": [0.08, 0.34, 0.45, 0.13], "legend": ["calm", "annoyed", "frustrated", "angry"], "confidence": 0.45}},
 "usage": {"input_tokens": 412, "output_tokens": 0}}

Decision Index engine

With the decision-index package installed in the same environment, the model runs in-process with the same scoring: python -m decision_index run --engine matilda_jev.engine:MatildaJevEngine --option checkpoint=checkpoints/matilda-jev ...

Limitations

  • Hard knowledge and multi-step reasoning are the weakest area (GPQA Diamond, MMLU-Pro, BBH): a single forward pass cannot reason at length the way a thinking model does.
  • Decisions only. It does not generate text or explanations; it scores the options you give it.
  • English. Decision training is in English. Images are accepted, but decision training was text-only.
  • Calibration was fitted on held-out general decision data; recalibrate on your own data for a new domain.
  • Not a substitute for review in high-stakes use (legal, medical, financial or safety decisions).

Licence and attribution

Apache 2.0

Citation

@misc{maincode2026matildajev,
  title  = {Matilda Jev: an Australian decision model},
  author = {Maincode},
  year   = {2026},
  url    = {https://huggingface.co/Maincode/matilda-jev}
}