Laya Mini CPU — research checkpoints and controlled A/B results

English decision models, approximately 87.16 MB each. Research artifacts, not a validated general-purpose decision agent.

Français · Benchmarks · Research notes · Reproducibility · Rights and restrictions

Maintained by VynoDePal. Research snapshot: 26 September 2026. This is an independent student-model experiment using Laya as teacher, not an official Laya release or an affiliation claim.

Read this before choosing a checkpoint

Experimental B improves accuracy on the tested close reformulations, but fails the preregistered composite acceptance criterion. It is NOT promoted as a replacement for the original. None of these variants is validated for reliable general decisions. The reserved FINAL set was not evaluated.

The main experiment covers 312 texts / 2,184 correlated requests, two domains, two description styles and one training seed. It is not an official full-dataset AG News or SNLI leaderboard evaluation.

DEV domain / precision Control A accuracy Experimental B accuracy A order instability B order instability Accuracy + stability gate
AG News FP32 57.05% 89.02% 61.54% 3.53% Pass
AG News INT8 54.89% 88.46% 65.06% 6.41% Pass
SNLI FP32 53.42% 71.05% 48.40% 16.03% Pass
SNLI INT8 50.11% 65.60% 58.01% 33.65% FAIL: B must be ≤29.01% unstable
  • The SNLI INT8 short-style gain is only 5.98 percentage points, not the 15.49-point average across styles.
  • B loses 5.45 points on SNLI from FP32 to INT8, despite fewer FP32/INT8 prediction changes than A. Fewer changes do not establish better quantization robustness.
  • Confidence intervals resample text/premise families, not individual rotations. They do not measure variation across training seeds.
  • experiment-status.json records quality_validated: false, dev_passed: false, the failed gate and final_evaluated: false.

Model variants

Directory Role CPU model payload FP32 checkpoint
models/original/ Historical reference, before the A/B continuation 87,157,026 bytes 152,909,884 bytes
models/control-a/ Two more epochs using fixed descriptions 87,155,068 bytes 152,909,884 bytes
models/experimental-b/ Same continuation, with controlled description augmentation 87,155,068 bytes 152,909,884 bytes

Payload means ONNX INT8 + tokenizer + manifest; 1 MB = 1,000,000 bytes. Each variant satisfies the 80–90 MB model-file target. The whole repository is larger because it contains multiple variants, optional FP32 reference checkpoints, code and evidence. Python, dependencies and working RAM are extra.

There is deliberately no default model at the repository root. Choose a variant explicitly. Original is a historical reference, not a claim that it is generally reliable. Checkpoints under checkpoints/ are for inspection, comparison and possible re-export; downloading them is unnecessary for CPU INT8 inference.

Download only the CPU variant you need

Install the Hugging Face CLI if needed (python -m pip install huggingface_hub), then:

hf download VynoDePal/laya-mini-cpu-research \
  --include 'models/experimental-b/*' runtime.py schema.py requirements-runtime.txt 'examples/*' \
  --local-dir laya-mini
cd laya-mini
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-runtime.txt
python runtime.py models/experimental-b examples/example-canonical-request.json --threads 2

Use models/original or models/control-a to select another variant. The custom runtime uses only CPUExecutionProvider; it does not require PyTorch, a GPU, paid inference or network access after installation. It is not a Transformers pipeline() checkpoint, an Ollama/GGUF model or a hosted chat endpoint.

import json
from pathlib import Path
from runtime import DecisionSession

session = DecisionSession("models/experimental-b", threads=2)
request = json.loads(Path("examples/example-canonical-request.json").read_text())
result = session.predict(request["state"], request["questions"])
print(result["answers"])

Keep the session loaded between requests. Inputs, instructions and option descriptions should be in English. The two included tennis examples are previously known interface checks, not held-out evidence. B answers sports on both in the recorded smoke test; the original fails the long-description version. Do not infer general reliability from either anecdote.

Interface and limits

  • choice: chooses a key from a dictionary of option descriptions.
  • score: returns the expectation over an ordered list of levels.
  • noul: returns a soft boolean score.
  • Software supports 2–20 options; training covered only 2–6.
  • Maximum total context: 512 WordPiece tokens. Instructions/options have a 192-token budget; oversized question definitions are rejected. Long states may be truncated; check truncated_state.
  • Probabilities are not calibrated. A high max_probability is not a factual correctness guarantee.
  • No free-form generation, vision, screen reading or automatic action execution. A game-playing system would still need an adapter and game-specific evaluation.

CPU performance

Intel Core i7-1260P, Linux x86-64, ONNX Runtime 1.24.4, two threads, one question / four options, warm model, 30 repetitions. Synthetic inputs; excludes tokenization and loading.

B input tokens Median latency
64 11.08 ms
128 21.55 ms
256 44.83 ms
512 111.79 ms

B's 128-token p95 is 28.83 ms, and peak process RSS is 196 MiB across the benchmark. Original / A / B 128-token medians measured in the same session are 21.36 / 21.43 / 21.55 ms. No meaningful speed ranking is established. The original's older 10.79 ms result came from a different session and is not a matched comparison; the difference is unexplained.

Raw per-repetition timings were not retained, so an IQR is unavailable. Windows, ARM64 and old/low-end CPUs were not validated. “Runs easily on every CPU” is not established. More settings and p95 values are in BENCHMARKS.md.

Architecture and training

  • 38,223,873 parameters: a pretrained six-layer, 512-wide BERT backbone, an additional decision layer, question-type embeddings and an option-marker scoring head.
  • Backbone: google/bert_uncased_L-6_H-512_A-8, revision dd53ec6ca9d05e0a91b309c4e137f31988888071.
  • Teacher: convaiinnovations/laya, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, package laya==0.3.20.
  • Loss: teacher KL at temperature 1 + 2 × cross-entropy against actual labels. This is distillation plus supervised learning, not direct quantization of the original Laya weights.
  • A/B both continue the same original checkpoint for exactly two epochs / 5,048 steps, seed 42, batch 32, identical schedules, shared padding and last-checkpoint selection. There is no DEV-based checkpoint selection.
  • B replaces exactly 12,000 AG News and 12,000 SNLI descriptions, 50% of each training pool. All other fields/tasks remain unchanged. Both arms share teacher targets on fixed prompts, not teacher re-annotations of the variants.
  • Student: CUDA BF16 with FP32 matmul mode high (TF32 allowed; kernels not traced). Shared teacher cache: FP32 CUDA without autocast/TF32. Bitwise training reproducibility is not established.
  • Export: ONNX opset 18; dynamic INT8 linear-weight quantization, per channel, reduce_range=True; embeddings remain FP32. The reduced range uses seven effective weight bits in INT8 storage. No separate calibration corpus was used.

The 80,758 training rows comprise 24,000 AG News, 24,000 SNLI, 15,966 Emotion, 8,533 SST5 and 8,259 BoolQ examples. The original model's development history and different historical evaluations are documented separately; old test/COPA scores must not be attributed to A or B.

Evidence and reproducibility boundary

  • A/B report: all acceptance criteria, per-style diagnostics and confidence intervals.
  • Frozen protocol: hyperparameters and cryptographic commitments to inputs, including unevaluated FINAL files.
  • Text-free recorded decisions: 2,184 DEV + 2,000 fixed-validation rows, allowing aggregate statistics to be recomputed without rerunning inference.
  • CPU reports: measurements for all three variants in the same session.
  • Historical results: clearly scoped to the original or the pre-A/B audit.
  • Research sources, checkpoint metadata and SHA256SUMS: code/provenance and integrity evidence.

No raw training corpus, teacher-cache contents, FINAL requests/descriptions, credentials, account balances or private conversation transcripts are included. The final-set preparation script is omitted. Published hashes/code are not a cryptographic guarantee that nobody could reconstruct held-out public examples; the claim is that FINAL was not evaluated in this experiment.

Recomputing the recorded statistics is reproducible here. Training from scratch and rerunning the original DEV inputs are not fully self-contained, because their data/cache artifacts are not redistributed. FP32 checkpoints do not remove that limitation. See REPRODUCIBILITY.md.

Rights, safety and status

This public research snapshot is released at the owner's request with unresolved upstream rights explicitly disclosed. The other license tag is a pointer to source-specific restrictions, not a newly granted unified license. The Apache-2.0 model cards of the teacher/backbone do not by themselves clear the mixed training-data rights. Emotion is research/education only, AG News has unknown/non-commercial wording, and the pinned SST5 card declares no license. See LICENSE_NOTES.md before redistribution or commercial use.

Do not use these experimental models for consequential decisions or irreversible actions without independent validation and safeguards. The size target is met; reliable general decision intelligence is not.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VynoDePal/laya-mini-cpu-research

Finetuned
(7)
this model

Datasets used to train VynoDePal/laya-mini-cpu-research