D1A-E4B v0.5 · MLX 8-bit

JohnP1/d1a-e4b v0.5 for Apple Silicon. The build is the same as v0.4's:

  • linear layers and token embeddings 8-bit (group 64), per-layer embeddings 4-bit;
  • pointer head fp32;
  • calibration temperature 1.78 in d1a_config.json;
  • about 6 GB on disk, plus 1 GB of photo, voice and video encoders.

Only the first weight shard differs from v0.4: the per-layer embeddings and the encoders are byte-identical. The scores below were measured on this build.

D1A is a small open decision model. One document and a set of typed questions go in (yes/no, choice, score), and a calibrated probability for every option comes out, in one forward pass, with no generated text. It speaks the TypeSafe System One API (POST /v1/systemone), so the TypeSafe SDK and existing clients work against it.

v0.5 brings back Kev's harder skills on top of v0.4's pull-request labeling: hard decisions 55% → 71%, developer-tool decisions 64% → 70%, and gains on held-out transfer (+2.7), Japanese (+1.1), routing (+1.5) and long documents (+0.9). The PR labels hold: The change-type label improves (88% → 90%). It costs some pull-request labeling: severity drops 2 points, and blast radius on the owner's own repositories drops noticeably (see Limits). v0.4 stays available for PR labeling alone. See Results.

What's inside

A LoRA adapter (rank 16 on the attention and MLP projections: q, k, v, o, gate, up, down) and a small pointer head (256-dimensional) on google/gemma-4-E4B at revision 411aa17b. Each version continues training the previous one, so v0.5 contains every stage below:

Version Stage Training data Settings
v0.1 General decisions decision-v7 suite: ~12,500 records from public classification datasets plus generated policy and rule data 2 epochs, lr 5e-5
(v0.2) Kev's later skills dates and missing evidence (1,425), hard decisions, developer-tool decisions and long documents (hard-v1, devtools-v1, documents-v1 training partitions, states up to 4,096 tokens), decision-v7 replay 4,000 1 epoch, lr 2e-5
v0.2 Japanese and agent routing Japanese decisions from JGLUE (JNLI, JCommonsenseQA, JSTS; 6,000), agent-factory and model-routing questions (2,745), decision-v7 replay 3,000 1 epoch, lr 5e-5
v0.3 Pull-request labeling 3,500 English and 768 Japanese pull requests labeled with change type, blast radius and severity (closed PRs of NousResearch/hermes-agent, MIT; Japanese copies machine-translated; severity follows the labeling job's rules for docs-only and dependency PRs), replay: decision-v7 1,500, JGLUE 500, routing 300 1 epoch, lr 5e-5, documents up to 1,536 tokens
v0.4 Pull-request labeling, round 2 the 4,063 English PRs v0.3 did not see, its 2,396 rare-label PRs (blast radius, P0, P1, P4) again, 2,185 more PRs with a blast-radius label, the 768 Japanese PRs again; replay: decision-v7 1,500, JGLUE 500, routing 300 (11,540 records) from v0.3, 1 epoch, lr 5e-5, documents up to 1,536 tokens, 1,443 steps on one L4
v0.5 Kev's skills, all of them the whole hard-v1 (6,000) and devtools-v1 (5,320) training partitions, documents-v1 1,000, JGLUE 650, PR labels 1,350 (so the earlier skills are replayed, not forgotten), decision-v7 replay 1,000: 15,320 records from v0.4 through a 300-step pilot on 2,400 records of the same mix; 1 epoch, lr 2e-5, states up to 6,400 tokens, 1,915 steps on one A100

Calibration: a single temperature, T = 1.78, fitted on pooled held-out rows (decision-v7 calibration split and PR-labeling development set; calibration error 0.081 before, 0.015 after, out of fold), stored in d1a_config.json.

Results

Held-out data only. Both builds are MLX 8-bit, scored with d1a.eval.benchmark through D1A's quality gate (scripts/quality_gate.py --head-run ... --card-suites), except where noted:

Set v0.4 v0.5
Hard decisions (hard-v1, 1,083 questions) 55.2% 71.5%
Developer-tool decisions (devtools-v1, 1,074) 63.8% 69.8%
Long documents (documents-v1, 920) 86.6% 87.5%
Held-out transfer-v4 (764, sources never trained on) 68.9% 71.6%
General decisions (decision-v7, 1,468) 84.7% 85.3%
JGLUE development (1,500, Japanese) 81.0% 82.1%
Model routing: generic (270) / hand-labelled (45) 97.4% / 100% 98.9% / 100%
Agent factory development (640 questions) 90.8% 90.5%
PR labels, English (953 PRs)¹: change type / severity / blast radius 87.7% / 77.8% / 54.0% 89.7% / 75.6% / 58.3%
PR labels, Japanese (92 PRs)¹ 74.1% 74.6%
39 recent PRs from the owner's repositories, labels checked by hand²: type / blast radius / severity 86.8% / 87.2% / 65.6% 84.2% / 74.4% / 65.6%

¹ Scored on the bf16 PyTorch checkpoint with the serving context, so every PR counts. The 8-bit run scored only the 183 shorter test PRs: 82.2% → 85.0%. ² The PRs the owner's labeling job saw on 2026-10-06 and 07, labelled by the owner on the playground's review page, which showed v0.4's answer. On 98 earlier decisions of that job, with review-bot labels, blast radius was 70% (v0.4) against 63% (v0.5) and type 86% for both.

Against Kev-4B's model card (fp32, development splits; an approximate comparison): hard decisions 78.6, developer tools 73.9, long documents 89.1, transfer 81.7. v0.5 closes most of the gap on developer tools and long documents, and two thirds of it on hard decisions.

Calibration: v0.5's single temperature fits its calibration sets (0.015 out of fold). On sets outside them it is less sure than it is right: on agent factory, mean confidence 78% at 90.5% accuracy; on hard decisions, 4 points under. Its probabilities err on the cautious side.

Photos, voice and video

media/ holds Gemma 4 E4B's own vision and audio encoders (bf16, ~1 GB, unchanged from the base: the LoRA adapter only touches the text model; byte-identical to v0.4's). With them d1a.serving.serve answers questions about a photo, a voice clip or a video with this same model.

  • The endpoint is POST /v1/systemone/media: a /v1/systemone request plus {"media": {"type": "image" | "audio" | "video", "data": <base64>}}.
  • The encoders turn the media into soft tokens, which the 8-bit language model reads right after <state>.
  • A video is read as 16 timestamped frames; its sound is not used.
  • The encoders are fetched on the first such request; a plain download skips them.

No media in training (zero-shot). On the playground's photo, voice and video examples, v0.5 changes 1 of 28 answers from v0.4.

The questions it was trained for

PR labeling works best with these questions, worded exactly so, over a document of the form title / author / stats / body / files (one - status path +added/-deleted line per file):

  • type (choice): "Primary change type from files and body, not the title prefix." Options: type/bug, type/docs, type/feature, type/perf, type/refactor, type/security, type/test.
  • blast (choice): "How far a mistake in this PR spreads in production." Options: review:blast-contained (one module), -moderate (one subsystem), -broad (shared helper or config), -massive (auth, permissions, or all paths).
  • sev (choice): "How serious the problem this PR addresses is — not the risk of merging the diff as-is." Options P0 to P4.

The PR-labeler recipe has the exact wording, the document builder and the training and scoring scripts. The skills recipe builds v0.5's mix. Other questions about other documents work as in v0.2.

Run it

On Apple Silicon (the PyTorch checkpoint, for NVIDIA GPUs and CPUs, is JohnP1/d1a-e4b v0.5):

pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a"
python -m d1a.serving.serve --run JohnP1/d1a-e4b-mlx-q8@v0.5 --port 8009

or in-process: from d1a import D1A; D1A.load("JohnP1/d1a-e4b-mlx-q8@v0.5").decide(state, questions).

Limits

  • Blast radius on the owner's own repositories is worse than v0.4: 74% against 87% on 39 hand-checked PRs, and 63% against 70% on 98 earlier decisions. On the larger test set from another project it is better (58% against 54%). Its errors there are mostly moderate PRs called contained. For PR labeling alone, v0.4 is the better choice for now; a follow-up round on PR labeling is planned.
  • Severity is weakest at the urgent end and 2 points below v0.4 on the English test (75.6% against 77.8%).
  • One temperature for every question: outside its calibration sets v0.5 is under-confident (see Results).
  • About 1% slower than v0.4 on short requests (same size and architecture; measured interleaved). Long documents take the same time.
  • Still below Kev-4B on hard decisions (−7) and held-out transfer (−10).
  • The PR data comes from one large open-source project; label conventions follow its maintainers except where the labeling job's rules override them.
  • Questions are in English; the documents can be English or Japanese.

Versions

Tag What it adds
v0.1 general decisions (decision-v7, 2 epochs, calibrated)
v0.2 Kev's later skills, Japanese, agent routing
v0.3 pull-request labeling, English and Japanese
v0.4 pull-request labeling round 2: severity, Japanese, less biased blast radius
v0.5 Kev's harder skills back (hard, developer-tool and long-document decisions), PR labels held

Older tag names (v0.1-2epoch, v0.1.1-2epoch-calibrated, v0.2-hybrid) still work.

License and data

Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0). Japanese decision data derived from JGLUE by Yahoo Japan Corporation and Waseda University (CC BY-SA 4.0). Pull-request data from NousResearch/hermes-agent (MIT); the labeled PR dataset itself is private.

Downloads last month
67
Safetensors
Model size
7B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JohnP1/d1a-e4b-mlx-q8

Quantized
(1)
this model