Akiki MiniLM-L6-v2: Typed Decisions specialist, epoch 158
64.50% accuracy under the published specialist's scoring convention:
1,290 / 2,000 correct when the target is the highest-probability gold answer.
The original experiment scored 64.60% (1,292 / 2,000) against the dataset's
saved gold.label. These are different conventions: the saved label and gold
probability argmax disagree on 31 of the 2,000 questions, all tied gold maxima.
The published scorer picks the first probability key in a tie; our original
scorer uses the stored gold label. Both results are retained
and reproducible; the Hugging Face evaluation YAML uses the published convention.
This is the selected epoch-158 checkpoint from refinement at 1e-5 for both Akiki and the classifier head. All 22,713,216 MiniLM parameters stayed frozen; only the input affine W and b (147,840 parameters) and ten-option head (3,850 parameters) were trained. Total parameters: 22,864,906.
The released MiniLM specialist reports 60.65%, rounded to 60.7% on its model card. With the same scoring convention, Akiki improves accuracy by 3.85 percentage points. The benchmark's ModernBERT-base specialist reports 64.60%, which is 0.10 percentage points above our matched-convention result. Our KL and soft Brier are higher than both published references.
| Model | Accuracy ↑ | KL from gold ↓ | Soft Brier ↓ | Macro per-question ECE ↓ |
|---|---|---|---|---|
| Akiki + MiniLM, epoch 158 | 64.50% | 0.251983 | 0.139726 | 0.132747 |
| Released MiniLM specialist | 60.65% | 0.248679 | 0.135730 | 0.135824 |
| ModernBERT-base specialist, published only | 64.60% | 0.223 | 0.119 | 0.179 |
MiniLM predictions and metrics were reproduced locally; ModernBERT was not run. These are our measurements, not a benchmark-maintainer certification. References were checked on 2026-10-08. Training setups differ. ECE is a macro average across workflow/question groups in the matched MiniLM comparison; our original report instead pooled all decisions. Both ECE forms are included in the result files.
Typed Decisions benchmark submission
Organization repository:
Akiki-AI/Akiki-MiniLM-L6-v2.
The four benchmark metrics (accuracy, kl_from_gold, brier, ece) are in
.eval_results/typed-decisions.yaml, with
exact values and scoring conventions in
results/metrics-gold-argmax.json.
This is a specialist trained on the benchmark's train split, with 1,080
training cases and 120 validation cases. The separate test split contains 400
cases / 2,000 decisions. Each evaluation call passes the complete state and all
five questions to Specialist.predict(state, questions). Internally, the model
processes questions in microbatches of four and one. All MiniLM parameters stay
frozen; only the input affine W, b and ten-option classifier head were trained.
Long states use 256-token chunks without discarding state tokens. The full
training recipe and chunk aggregation are documented below and in
experiment.yaml.
Reproduce the four submitted metrics from the packaged checkpoint:
python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax
# Or verify the scores from saved predictions without loading the model:
python reproduce.py --score-only --split test --gold-convention distribution_argmax
These are self-reported results following the benchmark's submission instructions.
Run the packaged model
Everything needed to load this checkpoint offline is included: base weights,
tokenizer, configuration, affine and head weights, data, and Python scripts.
All model loading uses local_files_only=True; offline environment flags are set
before importing Transformers. Missing files cause an error. No model is fetched.
Use Python 3.12 and the dependencies in requirements.txt. The tested runtime
was PyTorch 2.13.0+cu130, Transformers 5.17.0, CUDA 13.0. Install those
Python dependencies in an appropriate environment if needed; no installation is
performed by the reproduction scripts.
From this folder:
# Recompute the published-convention metrics from saved predictions; standard library only.
python reproduce.py --score-only --split both --gold-convention distribution_argmax
# Recompute the original experiment's saved-label metrics (64.60% test accuracy).
python reproduce.py --score-only --split both --gold-convention saved_label
# Rerun all 2,000 test decisions with the published scoring convention.
python reproduce.py --device cuda:0 --gold-convention distribution_argmax
# Also run all 600 validation decisions (CPU is supported via --device cpu).
python reproduce.py --device cuda:0 --split both --gold-convention distribution_argmax
Outputs go to reproduced/. The check requires identical predicted labels and
probability/metric differences within 1e-5. Numerical details can vary across
hardware, runtimes and attention kernels; cross-platform bitwise identity is not
promised. A fresh run on the RTX PRO 6000 Blackwell matched every one of the
2,600 validation and test labels, including 64.60% saved-label / 64.50% gold-argmax test accuracy; the largest
probability difference from the original RTX 4090 run was below 0.000006.
See checks/rtx6000-inference-verification.json.
Minimal use in Python:
import json
from model import Specialist
model = Specialist(device="cuda:0")
with open("data/test.jsonl") as f:
case = json.loads(next(f))
result = model.predict(case["state"], case["questions"])
print(result["answers"])
The custom loader is required. Calling AutoModel.from_pretrained("base_model")
loads only the original MiniLM encoder and does not apply the trained affine/head.
The ten outputs are option slots, not vocabulary tokens or ten fixed semantic
classes. The input supplies each question's option descriptions. Unused slots are
masked before softmax. The output contains a probability for every valid option.
Results and scoring conventions
The table below preserves the original saved-label experiment. For direct
comparison with the released specialist, use results/metrics-gold-argmax.json:
test accuracy 64.50%, NLL 0.849345, hard Brier 0.487758, pooled ECE 0.082780,
macro per-question ECE 0.132747. KL, soft Brier and soft cross entropy are unchanged
because they already use the complete gold probability distribution.
| Metric | Validation (600 decisions) | Test (2,000 decisions) |
|---|---|---|
| Accuracy | 63.6667% | 64.6000% |
| Hard-label NLL / log loss | 0.811622 | 0.851018 |
| KL(gold distribution ‖ prediction) | 0.242419 | 0.251983 |
| Soft-target cross entropy | 0.971203 | 1.019342 |
| Soft-target Brier | 0.136050 | 0.139726 |
| Hard-label Brier | 0.466743 | 0.488339 |
| Top-label ECE, 15 bins | 0.064592 | 0.083780 |
Logs are natural (nats). Both Brier variants sum over classes, then average
questions; they do not divide by the number of classes. The benchmark-facing
brier value is the soft-target variant. KL compares the full normalized gold
distribution to the prediction. ECE uses 15 uniform right-closed confidence bins.
Score-question accuracy uses the most probable level, not a rounded expected
score. Boolean predictions use P(false) = 1 - P(true), matching the original
scorer. No fitted temperature or test-set decision threshold is applied.
results/metrics.json contains exact original saved-label values;
results/metrics-gold-argmax.json contains the comparable published-convention values. *-predictions.jsonl
contains the original case-level outputs, and *-decisions.jsonl gives individual
question distributions, correctness, entropy and top-two margins. Reports include
workflow/type breakdowns and calibration summaries. The Hugging Face evaluation
YAML is .eval_results/typed-decisions.yaml.
Reproduce the training experiment
experiment.yaml records the complete recipe and exact split sizes. The
checkpoint uses the public Typed Decisions specialist training split: 1,080
training cases / 5,400 decisions and 120 validation cases / 600 decisions,
split with seed 42. The held-out test set contains 400 cases / 2,000 decisions.
Cases are disjoint by ID and serialized state; checksums and audit are bundled.
The data is synthetic and its target distributions are teacher-derived.
Training starts with W = identity, b = zero and a seeded random 384-to-10 head.
The affine transforms word embeddings before native BERT position/type
embeddings and LayerNorm. The full MiniLM encoder, including embeddings, blocks,
norms and pooler, remains frozen. Dropout stays disabled during training. There
is no language-model generation head. The optimizer updates only four tensors:
interface.proj.weight, interface.proj.bias, head.weight, head.bias.
Long contexts are split into non-overlapping chunks with at most 256 tokens, repeating the complete question and options in every chunk. No state tokens or options are discarded. Each chunk receives masked mean pooling and L2 normalization; chunk vectors are averaged with state-token-count weights and normalized again. The resulting 384-dimensional vector feeds the classifier.
Training uses soft-target cross entropy, AdamW (betas 0.9/0.999, epsilon 1e-8, weight decay 0), effective batch 32, microbatch 4, gradient-norm clipping at 1, FP32, and deterministic sample/option permutations for every epoch.
- Initial stage: affine LR 1e-4, head LR 1e-3, seed 0. The original run ended at epoch 167 after 10 epochs without a lower validation loss; epoch 157 was selected. The portable script runs this observed 167-epoch budget and selects the lowest validation loss.
- Refinement: reload selected epoch 157, reset AdamW, set both LRs to 1e-5. Run at least 10 new epochs and stop after 10 consecutive epochs without a strictly lower validation loss. Epoch 158 was best; this stage stopped at 168 after 11 new epochs.
- A subsequent 1e-6 trial did not improve validation loss and left epoch 158 selected. It is not part of the recipe needed to produce these weights.
# Full two-stage training; writes to a new output directory.
python train.py --device cuda:0 --mode full --output training-output
# Faster: reproduce epoch 158 from the included epoch-157 weights, fresh optimizer.
python train.py --device cuda:0 --mode refine --max-new-epochs 1 --output training-check
# Rerun the complete refinement stopping rule from epoch 157.
python train.py --device cuda:0 --mode refine --output refinement-output
The bundled refinement script was verified by replaying all 169 optimizer steps
from epoch 157 on the RTX PRO 6000 and re-evaluating all 2,000 test questions:
64.60% saved-label accuracy (64.50% gold-argmax accuracy), with every predicted
label matching the original epoch 158. The
maximum probability difference was below 0.000006. Fresh identity-affine and
seeded-head initialization were also checked. The full 167-epoch initial stage
was not rerun while preparing this package. See checks/training-reproduction.json.
The training script intentionally requires a fresh stage output directory. It
writes weights and history each epoch; it is a compact reproduction script, not
the original restartable service runner. Original training implementation and
per-epoch histories are in source/ and provenance/ for audit. Those archived
scripts are reference material; the root-level scripts are standalone.
Checkpoint selection used validation soft cross entropy. Test scores were repeatedly inspected during development, so this is not a blind one-shot test. This is a specialist trained on Typed Decisions examples, not a zero-shot general reasoning result. Only one training seed was used. Equal reported accuracy does not establish statistical equivalence or superiority across other tasks.
Latency on the benchmark GPU
Measured on NVIDIA RTX PRO 6000 Blackwell Workstation Edition, FP32, TF32 disabled, PyTorch eager SDPA, batch one, after 50 warm-up passes. Synthetic single-pass timings use 500 repetitions, already-tokenized input on the GPU, and include Akiki + MiniLM + pooling + classifier + softmax. Wall-clock times synchronize CUDA. End-to-end timings also include input serialization, tokenization, chunking, device transfer and output materialization.
| Measurement | Mean | Median | p95 |
|---|---|---|---|
| Single pass, 64 tokens | 1.25 ms | 1.24 ms | 1.29 ms |
| Single pass, 128 tokens | 1.31 ms | 1.30 ms | 1.33 ms |
| Single pass, 256 tokens | 1.39 ms | 1.38 ms | 1.41 ms |
| One complete TD decision, 500 samples | 2.32 ms | 2.40 ms | 2.73 ms |
| Five-question TD case, all 400 cases | 6.52 ms | 6.70 ms | 8.41 ms |
These are warm in-process timings, excluding model loading and network/server overhead. Long decisions can involve several context chunks; a five-question case uses two question microbatches (4 + 1). The inference benchmark's peak PyTorch allocated memory was about 192 MiB. CPU: Ryzen 9 9900X, four PyTorch CPU threads. These numbers are hardware/runtime specific.
python benchmark_latency.py --device cuda:0 --repeats 500 --warmup 50
Exact timings, method, sample IDs and environment are in
results/latency-rtx6000.json.
Matched latency comparison with the published MiniLM specialist
Both models were rerun on the RTX PRO 6000, FP32, four CPU threads, with 50 warm-up
cases and three rounds over the same 400 held-out cases. Model order alternated
between rounds. The released specialist used its original predict.py and
adaptive-classifier 0.2.0; only loading was redirected to the existing local base
model with local_files_only=True. Its saved reference probabilities matched
within 0.000002, with all 100 reference labels matching. Full-test accuracy, KL,
soft Brier, cross entropy and macro ECE also matched the authors' report.
| Model | Five-question case median | Case p95 | Single-question median |
|---|---|---|---|
| Released MiniLM specialist | 10.47 ms | 11.90 ms | 2.11 ms |
| Akiki epoch 158 | 6.72 ms | 7.92 ms | 2.49 ms |
Akiki's complete-case path is 1.56x faster in this measurement. The published model is faster for individual questions. Akiki batches questions (4 + 1); the published predictor makes five sequential encoder calls plus prototype/head scoring. These compare the shipped prediction paths, not identical inference implementations. The published model truncates at 512 tokens (63 of 2,000 test questions); Akiki retains all context through chunking.
The benchmark's 22 ms M3 Max figure belongs to its earlier 58.7% MiniLM row, not
the released 60.65% checkpoint, so a hardware-only speedup cannot be inferred from
those two numbers. Raw comparison metadata is in
results/rtx6000-specialist-comparison.json. Model loading and network/server
overhead are excluded throughout.
Files, integrity and attribution
adapter.safetensors: selected epoch-158 affine and head (607,080 bytes).base_model/: original locally cached MiniLM weights, tokenizer and configs.checkpoints/epoch-157.safetensors: source for the reproducible refinement.model.py,reproduce.py,train.py,benchmark_latency.py: portable code.experiment.yaml,model_config.json: training recipe and checkpoint identity.data/: exact train/validation/test JSONL used here.results/,checks/,provenance/: predictions, scores, verification and audit.SHA256SUMS: hashes of every packaged file other than this hash list itself.
Verify file integrity on Linux with sha256sum -c SHA256SUMS.
All assets are local files, with no symlinks back to a cache or source workspace.
No uploads are performed by any script. See LICENSE, NOTICE, and the original
base model card in base_model/README.md for Apache-2.0 licensing and attribution.
Model tree for Akiki-AI/Akiki-MiniLM-L6-v2
Base model
nreimers/MiniLM-L6-H384-uncasedDataset used to train Akiki-AI/Akiki-MiniLM-L6-v2
Evaluation results
- LocalLLaMA/typed-decisions leaderboard
- Accuracy View evaluation resultsSelf-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Gold answer is the argmax of its probability distribution; first key wins ties. 1290/2000 correct (0.645); original saved-label accuracy is 0.646. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.0.65 *
- Kl From Gold View evaluation resultsSelf-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Mean KL(gold distribution || predicted distribution), using natural logarithms and full answer distributions. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.0.25 *
- Brier View evaluation resultsSelf-reported specialist trained on LocalLLaMA/typed-decisions train: 1080 training cases, 120 validation cases. Frozen MiniLM backbone; only input affine W,b and ten-option head trained. Selected epoch 158 by validation soft cross entropy. Test: 400 separate cases / 2000 decisions, one complete state-plus-five-questions call per case. Soft-target Brier: sum of squared differences from the gold probability distribution over valid options, averaged over decisions; no division by class count. Reproduce: python reproduce.py --device cuda:0 --split test --gold-convention distribution_argmax.0.14 *