You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Akiki MiniLM-L6-v2: Typed Decisions specialist, epoch 158

64.50% accuracy under the published specialist's scoring convention: 1,290 / 2,000 correct when the target is the highest-probability gold answer. The original experiment scored 64.60% (1,292 / 2,000) against the dataset's saved gold.label. These are different conventions: the saved label and gold probability argmax disagree on 31 of the 2,000 questions, all tied gold maxima. The published scorer picks the first probability key in a tie; our original scorer uses the stored gold label. Both results are retained and reproducible; the Hugging Face evaluation YAML uses the published convention.

This is the selected epoch-158 checkpoint from refinement at 1e-5 for both Akiki and the classifier head. All 22,713,216 MiniLM parameters stayed frozen; only the input affine W and b (147,840 parameters) and ten-option head (3,850 parameters) were trained. Total parameters: 22,864,906.

The released MiniLM specialist reports 60.65%, rounded to 60.7% on its model card. With the same scoring convention, Akiki improves accuracy by 3.85 percentage points. The benchmark's ModernBERT-base specialist reports 64.60%, which is 0.10 percentage points above our matched-convention result. Our KL and soft Brier are higher than both published references.

Model Accuracy ↑ KL from gold ↓ Soft Brier ↓ Macro per-question ECE ↓
Akiki + MiniLM, epoch 158 64.50% 0.251983 0.139726 0.132747
Released MiniLM specialist 60.65% 0.248679 0.135730 0.135824
ModernBERT-base specialist, published only 64.60% 0.223 0.119 0.179

MiniLM predictions and metrics were reproduced locally; ModernBERT was not run. These are our measurements, not a benchmark-maintainer certification. References were checked on 2026-10-08. Training setups differ. ECE is a macro average across workflow/question groups in the matched MiniLM comparison; our original report instead pooled all decisions. Both ECE forms are included in the result files.

Typed Decisions benchmark submission

Organization repository: Akiki-AI/Akiki-MiniLM-L6-v2. The four benchmark metrics (accuracy, kl_from_gold, brier, ece) are in .eval_results/typed-decisions.yaml, with exact values and scoring conventions in results/metrics-gold-argmax.json.

This is a specialist trained on the benchmark's train split, with 1,080 training cases and 120 validation cases. The separate test split contains 400 cases / 2,000 decisions. Each evaluation call passes the complete state and all five questions to Specialist.predict(state, questions). Internally, the model processes questions in microbatches of four and one. All MiniLM parameters stay frozen; only the input affine W, b and ten-option classifier head were trained. Long states use 256-token chunks without discarding state tokens. The full training recipe and chunk aggregation are documented below and in experiment.yaml.

Reproduce the four submitted metrics from the packaged checkpoint:

python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax
# Or verify the scores from saved predictions without loading the model:
python reproduce.py --score-only --split test --gold-convention distribution_argmax

These are self-reported results following the benchmark's submission instructions.

Run the packaged model

Everything needed to load this checkpoint offline is included: base weights, tokenizer, configuration, affine and head weights, data, and Python scripts. All model loading uses local_files_only=True; offline environment flags are set before importing Transformers. Missing files cause an error. No model is fetched.

Use Python 3.12 and the dependencies in requirements.txt. The tested runtime was PyTorch 2.13.0+cu130, Transformers 5.17.0, CUDA 13.0. Install those Python dependencies in an appropriate environment if needed; no installation is performed by the reproduction scripts.

From this folder:

# Recompute the published-convention metrics from saved predictions; standard library only.
python reproduce.py --score-only --split both --gold-convention distribution_argmax

# Recompute the original experiment's saved-label metrics (64.60% test accuracy).
python reproduce.py --score-only --split both --gold-convention saved_label

# Rerun all 2,000 test decisions with the published scoring convention.
python reproduce.py --device cuda:0 --gold-convention distribution_argmax

# Also run all 600 validation decisions (CPU is supported via --device cpu).
python reproduce.py --device cuda:0 --split both --gold-convention distribution_argmax

Outputs go to reproduced/. The check requires identical predicted labels and probability/metric differences within 1e-5. Numerical details can vary across hardware, runtimes and attention kernels; cross-platform bitwise identity is not promised. A fresh run on the RTX PRO 6000 Blackwell matched every one of the 2,600 validation and test labels, including 64.60% saved-label / 64.50% gold-argmax test accuracy; the largest probability difference from the original RTX 4090 run was below 0.000006. See checks/rtx6000-inference-verification.json.

Minimal use in Python:

import json
from model import Specialist

model = Specialist(device="cuda:0")
with open("data/test.jsonl") as f:
    case = json.loads(next(f))
result = model.predict(case["state"], case["questions"])
print(result["answers"])

The custom loader is required. Calling AutoModel.from_pretrained("base_model") loads only the original MiniLM encoder and does not apply the trained affine/head. The ten outputs are option slots, not vocabulary tokens or ten fixed semantic classes. The input supplies each question's option descriptions. Unused slots are masked before softmax. The output contains a probability for every valid option.

Results and scoring conventions

The table below preserves the original saved-label experiment. For direct comparison with the released specialist, use results/metrics-gold-argmax.json: test accuracy 64.50%, NLL 0.849345, hard Brier 0.487758, pooled ECE 0.082780, macro per-question ECE 0.132747. KL, soft Brier and soft cross entropy are unchanged because they already use the complete gold probability distribution.

Metric Validation (600 decisions) Test (2,000 decisions)
Accuracy 63.6667% 64.6000%
Hard-label NLL / log loss 0.811622 0.851018
KL(gold distribution ‖ prediction) 0.242419 0.251983
Soft-target cross entropy 0.971203 1.019342
Soft-target Brier 0.136050 0.139726
Hard-label Brier 0.466743 0.488339
Top-label ECE, 15 bins 0.064592 0.083780

Logs are natural (nats). Both Brier variants sum over classes, then average questions; they do not divide by the number of classes. The benchmark-facing brier value is the soft-target variant. KL compares the full normalized gold distribution to the prediction. ECE uses 15 uniform right-closed confidence bins. Score-question accuracy uses the most probable level, not a rounded expected score. Boolean predictions use P(false) = 1 - P(true), matching the original scorer. No fitted temperature or test-set decision threshold is applied.

results/metrics.json contains exact original saved-label values; results/metrics-gold-argmax.json contains the comparable published-convention values. *-predictions.jsonl contains the original case-level outputs, and *-decisions.jsonl gives individual question distributions, correctness, entropy and top-two margins. Reports include workflow/type breakdowns and calibration summaries. The Hugging Face evaluation YAML is .eval_results/typed-decisions.yaml.

Reproduce the training experiment

experiment.yaml records the complete recipe and exact split sizes. The checkpoint uses the public Typed Decisions specialist training split: 1,080 training cases / 5,400 decisions and 120 validation cases / 600 decisions, split with seed 42. The held-out test set contains 400 cases / 2,000 decisions. Cases are disjoint by ID and serialized state; checksums and audit are bundled. The data is synthetic and its target distributions are teacher-derived.

Training starts with W = identity, b = zero and a seeded random 384-to-10 head. The affine transforms word embeddings before native BERT position/type embeddings and LayerNorm. The full MiniLM encoder, including embeddings, blocks, norms and pooler, remains frozen. Dropout stays disabled during training. There is no language-model generation head. The optimizer updates only four tensors: interface.proj.weight, interface.proj.bias, head.weight, head.bias.

Long contexts are split into non-overlapping chunks with at most 256 tokens, repeating the complete question and options in every chunk. No state tokens or options are discarded. Each chunk receives masked mean pooling and L2 normalization; chunk vectors are averaged with state-token-count weights and normalized again. The resulting 384-dimensional vector feeds the classifier.

Training uses soft-target cross entropy, AdamW (betas 0.9/0.999, epsilon 1e-8, weight decay 0), effective batch 32, microbatch 4, gradient-norm clipping at 1, FP32, and deterministic sample/option permutations for every epoch.

  1. Initial stage: affine LR 1e-4, head LR 1e-3, seed 0. The original run ended at epoch 167 after 10 epochs without a lower validation loss; epoch 157 was selected. The portable script runs this observed 167-epoch budget and selects the lowest validation loss.
  2. Refinement: reload selected epoch 157, reset AdamW, set both LRs to 1e-5. Run at least 10 new epochs and stop after 10 consecutive epochs without a strictly lower validation loss. Epoch 158 was best; this stage stopped at 168 after 11 new epochs.
  3. A subsequent 1e-6 trial did not improve validation loss and left epoch 158 selected. It is not part of the recipe needed to produce these weights.
# Full two-stage training; writes to a new output directory.
python train.py --device cuda:0 --mode full --output training-output

# Faster: reproduce epoch 158 from the included epoch-157 weights, fresh optimizer.
python train.py --device cuda:0 --mode refine --max-new-epochs 1 --output training-check

# Rerun the complete refinement stopping rule from epoch 157.
python train.py --device cuda:0 --mode refine --output refinement-output

The bundled refinement script was verified by replaying all 169 optimizer steps from epoch 157 on the RTX PRO 6000 and re-evaluating all 2,000 test questions: 64.60% saved-label accuracy (64.50% gold-argmax accuracy), with every predicted label matching the original epoch 158. The maximum probability difference was below 0.000006. Fresh identity-affine and seeded-head initialization were also checked. The full 167-epoch initial stage was not rerun while preparing this package. See checks/training-reproduction.json.

The training script intentionally requires a fresh stage output directory. It writes weights and history each epoch; it is a compact reproduction script, not the original restartable service runner. Original training implementation and per-epoch histories are in source/ and provenance/ for audit. Those archived scripts are reference material; the root-level scripts are standalone.

Checkpoint selection used validation soft cross entropy. Test scores were repeatedly inspected during development, so this is not a blind one-shot test. This is a specialist trained on Typed Decisions examples, not a zero-shot general reasoning result. Only one training seed was used. Equal reported accuracy does not establish statistical equivalence or superiority across other tasks.

Latency on the benchmark GPU

Measured on NVIDIA RTX PRO 6000 Blackwell Workstation Edition, FP32, TF32 disabled, PyTorch eager SDPA, batch one, after 50 warm-up passes. Synthetic single-pass timings use 500 repetitions, already-tokenized input on the GPU, and include Akiki + MiniLM + pooling + classifier + softmax. Wall-clock times synchronize CUDA. End-to-end timings also include input serialization, tokenization, chunking, device transfer and output materialization.

Measurement Mean Median p95
Single pass, 64 tokens 1.25 ms 1.24 ms 1.29 ms
Single pass, 128 tokens 1.31 ms 1.30 ms 1.33 ms
Single pass, 256 tokens 1.39 ms 1.38 ms 1.41 ms
One complete TD decision, 500 samples 2.32 ms 2.40 ms 2.73 ms
Five-question TD case, all 400 cases 6.52 ms 6.70 ms 8.41 ms

These are warm in-process timings, excluding model loading and network/server overhead. Long decisions can involve several context chunks; a five-question case uses two question microbatches (4 + 1). The inference benchmark's peak PyTorch allocated memory was about 192 MiB. CPU: Ryzen 9 9900X, four PyTorch CPU threads. These numbers are hardware/runtime specific.

python benchmark_latency.py --device cuda:0 --repeats 500 --warmup 50

Exact timings, method, sample IDs and environment are in results/latency-rtx6000.json.

Matched latency comparison with the published MiniLM specialist

Both models were rerun on the RTX PRO 6000, FP32, four CPU threads, with 50 warm-up cases and three rounds over the same 400 held-out cases. Model order alternated between rounds. The released specialist used its original predict.py and adaptive-classifier 0.2.0; only loading was redirected to the existing local base model with local_files_only=True. Its saved reference probabilities matched within 0.000002, with all 100 reference labels matching. Full-test accuracy, KL, soft Brier, cross entropy and macro ECE also matched the authors' report.

Model Five-question case median Case p95 Single-question median
Released MiniLM specialist 10.47 ms 11.90 ms 2.11 ms
Akiki epoch 158 6.72 ms 7.92 ms 2.49 ms

Akiki's complete-case path is 1.56x faster in this measurement. The published model is faster for individual questions. Akiki batches questions (4 + 1); the published predictor makes five sequential encoder calls plus prototype/head scoring. These compare the shipped prediction paths, not identical inference implementations. The published model truncates at 512 tokens (63 of 2,000 test questions); Akiki retains all context through chunking.

The benchmark's 22 ms M3 Max figure belongs to its earlier 58.7% MiniLM row, not the released 60.65% checkpoint, so a hardware-only speedup cannot be inferred from those two numbers. Raw comparison metadata is in results/rtx6000-specialist-comparison.json. Model loading and network/server overhead are excluded throughout.

Files, integrity and attribution

  • adapter.safetensors: selected epoch-158 affine and head (607,080 bytes).
  • base_model/: original locally cached MiniLM weights, tokenizer and configs.
  • checkpoints/epoch-157.safetensors: source for the reproducible refinement.
  • model.py, reproduce.py, train.py, benchmark_latency.py: portable code.
  • experiment.yaml, model_config.json: training recipe and checkpoint identity.
  • data/: exact train/validation/test JSONL used here.
  • results/, checks/, provenance/: predictions, scores, verification and audit.
  • SHA256SUMS: hashes of every packaged file other than this hash list itself.

Verify file integrity on Linux with sha256sum -c SHA256SUMS. All assets are local files, with no symlinks back to a cache or source workspace. No uploads are performed by any script. See LICENSE, NOTICE, and the original base model card in base_model/README.md for Apache-2.0 licensing and attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Akiki-AI/Akiki-MiniLM-L6-v2

Dataset used to train Akiki-AI/Akiki-MiniLM-L6-v2

Evaluation results

  • LocalLLaMA/typed-decisions leaderboard
  • Accuracy View evaluation results
    Self-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Gold answer is the argmax of its probability distribution; first key wins ties. 1290/2000 correct (0.645); original saved-label accuracy is 0.646. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.
    0.65 *
  • Kl From Gold View evaluation results
    Self-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Mean KL(gold distribution || predicted distribution), using natural logarithms and full answer distributions. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.
    0.25 *
  • Brier View evaluation results
    Self-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Soft-target Brier: sum of squared differences from the gold probability distribution over valid options, averaged over decisions; no division by class count. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.
    0.14 *