--- pretty_name: Matilda Jev license: apache-2.0 license_name: matilda-jev language: - en library_name: transformers tags: - decision-model - typed-decisions - jev - decision-index - maincode - australia --- # Matilda Jev > **TL;DR.** Matilda Jev is Australia's first Jev-style > decision model, conditioned for Australian compliance and political > even-handedness, then post-trained with our own specialised recipe. Scored > with the official Jev Decision Index kit, it reaches a score of **59.26**, > beating Jev which scored **57.91**. Much of the work people hand to language models is not open-ended writing. It is decisions: which tool to call, which label applies, which document answers the question, whether a claim is supported. Jev (Typesafe.ai) showed that these are best served by a model that answers them directly, as probabilities, rather than by parsing generated text. Matilda Jev is Maincode's model of that kind, built so that it behaves like a model deployed in Australia without giving up capability. | | | |---|---| | **Developer** | [Maincode](https://maincode.com), Australia | | **Head** | Same as Jev, 255-way answer-code readout over the backbone's final hidden state (1.3M parameters) | | **Output** | A probability for every option of a `choice`, `noul` (yes/no) or `score` question, temperature-calibrated | | **Precision / size** | bf16, ~49 GiB of weights | | **Context** | Up to 131k tokens per question on GPUs with ≥ 200 GB, 32k on 96 GB cards (the server sizes this automatically) | | **Inputs** | Text or JSON state; optional images (up to 4) | | **Languages** | English | ## How it works Each question is rendered into one prompt (state, question, lettered options, "return only the letter code") with thinking disabled. The model reads it once. A readout scores the answer codes `A`, `B`, … (up to 255 options, each a single token), and a stored temperature turns those scores into calibrated probabilities. There is no sampling and no answer parsing, so the output is deterministic and every option always gets a probability. - `choice`: pick among named options (1–255). Returns the distribution and the most likely option. - `noul`: a yes/no question. Returns P(yes). - `score`: an ordered rubric (2–10 levels). Returns the distribution and its expected level. A request can carry several questions over one shared state; each is one forward pass, batched together. ## Results Decision Index 0.2.1, scored with the official kit and scorer over the full suite (150,317 requests; every request answered, none truncated). Chance-corrected: 0 = random guessing, 100 = perfect. Board entries as published on the [leaderboard](https://huggingface.co/spaces/multimodalart/jev-decision-index) on 2026-09-28. | Model | Params | Decision Index | Raw | Breadth | |---|---|---:|---:|---:| | **Matilda Jev** | 26.1B | **59.26** | **68.89** | **58.06** | | Jev (Typesafe.ai, hosted) | – | 57.91 | 68.09 | 57.08 | | Surogate Rune 26B-A4B v3 | 25.8B | 57.44 | 67.30 | 56.47 | | Decider chat | 32.7B | 57.33 | 67.22 | 56.20 | | AutoJev-27B | 27.8B | 56.40 | 66.89 | 54.93 | | simple-jev | 27.8B | 55.74 | 66.26 | 53.91 | | Jebadiah 27B | 27.8B | 54.67 | 65.62 | 53.18 | | Eikos-27B | 27.8B | 53.13 | 63.64 | 52.02 | **By area** (chance-corrected skill × 100): | Model | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Human Taste | |---|---:|---:|---:|---:|---:| | **Matilda Jev** | **43.52** | **66.36** | **61.34** | **77.72** | **43.65** | | Jev | 51.40 | 62.02 | 55.42 | 75.09 | 37.66 | | Surogate Rune 26B-A4B v3 | 43.36 | 63.06 | 63.52 | 71.22 | 41.90 | | AutoJev-27B | 40.88 | 63.46 | 54.86 | 79.35 | 39.39 | **Latency.** Median 56.6 ms per request across the suite on 1× AMD Instinct MI355X, bf16, one request at a time. The board re-measures latency on its own RTX PRO 6000. **Contamination.** Every training row whose text copies a Decision Index test item was removed before training (exact-copy check against all 155,390 suite requests: 0 remaining). We are submitting the model to the board so the score can be verified with its own hardware, contamination and option-shuffle checks. **Hardware.** Trained on one node of 8× AMD Instinct MI355X (288 GB each) in Maincode's MC-2 cluster in Australia, with full-weight bf16 training. ## Run Python 3.12+, [uv](https://docs.astral.sh/uv/), and a GPU with room for ~49 GiB of bf16 weights plus runtime overhead (80 GB or more recommended; tested on AMD MI355X, targeted at NVIDIA RTX PRO 6000). ```bash git clone https://github.com/Maincode/matilda-jev.git cd matilda-jev uv sync --frozen --python 3.12 --extra cuda # NVIDIA (driver >= 580); use --extra rocm on AMD uv run --no-sync hf download Maincode/matilda-jev --local-dir checkpoints/matilda-jev MJ_CHECKPOINT=checkpoints/matilda-jev uv run --no-sync matilda-jev-serve ``` The API is at (interactive docs at `/docs`). `POST /v1/systemone` supports `choice`, `noul`, `score`, and optional base64 images. Set `MJ_API_KEY` (or `MJ_API_KEY_FILE`) to enable authentication and `MJ_HOST=0.0.0.0` to listen beyond localhost. While the weights are private, authenticate with `uv run --no-sync hf auth login` before downloading. Alternatives: `scripts/serve.sh [--background]` picks the CUDA or ROCm stack for you and reads `serve.env`; `docker compose up --build` runs it on NVIDIA with `./checkpoints` mounted read-only. ### Request ```bash curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{ "model": "matilda-jev-latest", "state": {"ticket": "My card was charged twice for the same order."}, "questions": { "intent": {"type": "choice", "instructions": "Main issue?", "criteria": {"duplicate_charge": "Charged twice for one purchase", "card_delivery": null, "other": null}}, "urgent": {"type": "noul", "instructions": "Does this need a same-day response?"}, "anger": {"type": "score", "instructions": "How frustrated is the customer?", "criteria": ["calm", "annoyed", "frustrated", "angry"]} } }' ``` ```json {"model": "matilda-jev", "answers": { "intent": {"type": "choice", "choice": "duplicate_charge", "probabilities": {"duplicate_charge": 0.97, "card_delivery": 0.01, "other": 0.02}, "confidence": 0.97}, "urgent": {"type": "noul", "noul": 0.81}, "anger": {"type": "score", "score": 1.6, "probabilities": [0.08, 0.34, 0.45, 0.13], "legend": ["calm", "annoyed", "frustrated", "angry"], "confidence": 0.45}}, "usage": {"input_tokens": 412, "output_tokens": 0}} ``` ### Decision Index engine With the `decision-index` package installed in the same environment, the model runs in-process with the same scoring: `python -m decision_index run --engine matilda_jev.engine:MatildaJevEngine --option checkpoint=checkpoints/matilda-jev ...` ## Limitations - **Hard knowledge and multi-step reasoning** are the weakest area (GPQA Diamond, MMLU-Pro, BBH): a single forward pass cannot reason at length the way a thinking model does. - **Decisions only.** It does not generate text or explanations; it scores the options you give it. - **English.** Decision training is in English. Images are accepted, but decision training was text-only. - **Calibration** was fitted on held-out general decision data; recalibrate on your own data for a new domain. - **Not a substitute for review** in high-stakes use (legal, medical, financial or safety decisions). ## Licence and attribution Apache 2.0 ## Citation ```bibtex @misc{maincode2026matildajev, title = {Matilda Jev: an Australian decision model}, author = {Maincode}, year = {2026}, url = {https://huggingface.co/Maincode/matilda-jev} } ```