Laya Mini CPU — research checkpoints and controlled A/B results
English decision models, approximately 87.16 MB each. Research artifacts, not a validated general-purpose decision agent.
Français · Benchmarks · Research notes · Reproducibility · Rights and restrictions
Maintained by VynoDePal. Research snapshot: 26 September 2026. This is an independent student-model experiment using Laya as teacher, not an official Laya release or an affiliation claim.
Read this before choosing a checkpoint
Experimental B improves accuracy on the tested close reformulations, but fails the preregistered composite acceptance criterion. It is NOT promoted as a replacement for the original. None of these variants is validated for reliable general decisions. The reserved FINAL set was not evaluated.
The main experiment covers 312 texts / 2,184 correlated requests, two domains, two description styles and one training seed. It is not an official full-dataset AG News or SNLI leaderboard evaluation.
| DEV domain / precision | Control A accuracy | Experimental B accuracy | A order instability | B order instability | Accuracy + stability gate |
|---|---|---|---|---|---|
| AG News FP32 | 57.05% | 89.02% | 61.54% | 3.53% | Pass |
| AG News INT8 | 54.89% | 88.46% | 65.06% | 6.41% | Pass |
| SNLI FP32 | 53.42% | 71.05% | 48.40% | 16.03% | Pass |
| SNLI INT8 | 50.11% | 65.60% | 58.01% | 33.65% | FAIL: B must be ≤29.01% unstable |
- The SNLI INT8 short-style gain is only 5.98 percentage points, not the 15.49-point average across styles.
- B loses 5.45 points on SNLI from FP32 to INT8, despite fewer FP32/INT8 prediction changes than A. Fewer changes do not establish better quantization robustness.
- Confidence intervals resample text/premise families, not individual rotations. They do not measure variation across training seeds.
experiment-status.jsonrecordsquality_validated: false,dev_passed: false, the failed gate andfinal_evaluated: false.
Model variants
| Directory | Role | CPU model payload | FP32 checkpoint |
|---|---|---|---|
models/original/ |
Historical reference, before the A/B continuation | 87,157,026 bytes | 152,909,884 bytes |
models/control-a/ |
Two more epochs using fixed descriptions | 87,155,068 bytes | 152,909,884 bytes |
models/experimental-b/ |
Same continuation, with controlled description augmentation | 87,155,068 bytes | 152,909,884 bytes |
Payload means ONNX INT8 + tokenizer + manifest; 1 MB = 1,000,000 bytes. Each variant satisfies the 80–90 MB model-file target. The whole repository is larger because it contains multiple variants, optional FP32 reference checkpoints, code and evidence. Python, dependencies and working RAM are extra.
There is deliberately no default model at the repository root. Choose a variant explicitly. Original is a historical reference, not a claim that it is generally reliable. Checkpoints under checkpoints/ are for inspection, comparison and possible re-export; downloading them is unnecessary for CPU INT8 inference.
Download only the CPU variant you need
Install the Hugging Face CLI if needed (python -m pip install huggingface_hub), then:
hf download VynoDePal/laya-mini-cpu-research \
--include 'models/experimental-b/*' runtime.py schema.py requirements-runtime.txt 'examples/*' \
--local-dir laya-mini
cd laya-mini
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-runtime.txt
python runtime.py models/experimental-b examples/example-canonical-request.json --threads 2
Use models/original or models/control-a to select another variant. The custom runtime uses only CPUExecutionProvider; it does not require PyTorch, a GPU, paid inference or network access after installation. It is not a Transformers pipeline() checkpoint, an Ollama/GGUF model or a hosted chat endpoint.
import json
from pathlib import Path
from runtime import DecisionSession
session = DecisionSession("models/experimental-b", threads=2)
request = json.loads(Path("examples/example-canonical-request.json").read_text())
result = session.predict(request["state"], request["questions"])
print(result["answers"])
Keep the session loaded between requests. Inputs, instructions and option descriptions should be in English. The two included tennis examples are previously known interface checks, not held-out evidence. B answers sports on both in the recorded smoke test; the original fails the long-description version. Do not infer general reliability from either anecdote.
Interface and limits
choice: chooses a key from a dictionary of option descriptions.score: returns the expectation over an ordered list of levels.noul: returns a soft boolean score.- Software supports 2–20 options; training covered only 2–6.
- Maximum total context: 512 WordPiece tokens. Instructions/options have a 192-token budget; oversized question definitions are rejected. Long states may be truncated; check
truncated_state. - Probabilities are not calibrated. A high
max_probabilityis not a factual correctness guarantee. - No free-form generation, vision, screen reading or automatic action execution. A game-playing system would still need an adapter and game-specific evaluation.
CPU performance
Intel Core i7-1260P, Linux x86-64, ONNX Runtime 1.24.4, two threads, one question / four options, warm model, 30 repetitions. Synthetic inputs; excludes tokenization and loading.
| B input tokens | Median latency |
|---|---|
| 64 | 11.08 ms |
| 128 | 21.55 ms |
| 256 | 44.83 ms |
| 512 | 111.79 ms |
B's 128-token p95 is 28.83 ms, and peak process RSS is 196 MiB across the benchmark. Original / A / B 128-token medians measured in the same session are 21.36 / 21.43 / 21.55 ms. No meaningful speed ranking is established. The original's older 10.79 ms result came from a different session and is not a matched comparison; the difference is unexplained.
Raw per-repetition timings were not retained, so an IQR is unavailable. Windows, ARM64 and old/low-end CPUs were not validated. “Runs easily on every CPU” is not established. More settings and p95 values are in BENCHMARKS.md.
Architecture and training
- 38,223,873 parameters: a pretrained six-layer, 512-wide BERT backbone, an additional decision layer, question-type embeddings and an option-marker scoring head.
- Backbone:
google/bert_uncased_L-6_H-512_A-8, revisiondd53ec6ca9d05e0a91b309c4e137f31988888071. - Teacher:
convaiinnovations/laya, revision55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, packagelaya==0.3.20. - Loss: teacher KL at temperature 1 + 2 × cross-entropy against actual labels. This is distillation plus supervised learning, not direct quantization of the original Laya weights.
- A/B both continue the same original checkpoint for exactly two epochs / 5,048 steps, seed 42, batch 32, identical schedules, shared padding and last-checkpoint selection. There is no DEV-based checkpoint selection.
- B replaces exactly 12,000 AG News and 12,000 SNLI descriptions, 50% of each training pool. All other fields/tasks remain unchanged. Both arms share teacher targets on fixed prompts, not teacher re-annotations of the variants.
- Student: CUDA BF16 with FP32 matmul mode
high(TF32 allowed; kernels not traced). Shared teacher cache: FP32 CUDA without autocast/TF32. Bitwise training reproducibility is not established. - Export: ONNX opset 18; dynamic INT8 linear-weight quantization, per channel,
reduce_range=True; embeddings remain FP32. The reduced range uses seven effective weight bits in INT8 storage. No separate calibration corpus was used.
The 80,758 training rows comprise 24,000 AG News, 24,000 SNLI, 15,966 Emotion, 8,533 SST5 and 8,259 BoolQ examples. The original model's development history and different historical evaluations are documented separately; old test/COPA scores must not be attributed to A or B.
Evidence and reproducibility boundary
- A/B report: all acceptance criteria, per-style diagnostics and confidence intervals.
- Frozen protocol: hyperparameters and cryptographic commitments to inputs, including unevaluated FINAL files.
- Text-free recorded decisions: 2,184 DEV + 2,000 fixed-validation rows, allowing aggregate statistics to be recomputed without rerunning inference.
- CPU reports: measurements for all three variants in the same session.
- Historical results: clearly scoped to the original or the pre-A/B audit.
- Research sources, checkpoint metadata and
SHA256SUMS: code/provenance and integrity evidence.
No raw training corpus, teacher-cache contents, FINAL requests/descriptions, credentials, account balances or private conversation transcripts are included. The final-set preparation script is omitted. Published hashes/code are not a cryptographic guarantee that nobody could reconstruct held-out public examples; the claim is that FINAL was not evaluated in this experiment.
Recomputing the recorded statistics is reproducible here. Training from scratch and rerunning the original DEV inputs are not fully self-contained, because their data/cache artifacts are not redistributed. FP32 checkpoints do not remove that limitation. See REPRODUCIBILITY.md.
Rights, safety and status
This public research snapshot is released at the owner's request with unresolved upstream rights explicitly disclosed. The other license tag is a pointer to source-specific restrictions, not a newly granted unified license. The Apache-2.0 model cards of the teacher/backbone do not by themselves clear the mixed training-data rights. Emotion is research/education only, AG News has unknown/non-commercial wording, and the pinned SST5 card declares no license. See LICENSE_NOTES.md before redistribution or commercial use.
Do not use these experimental models for consequential decisions or irreversible actions without independent validation and safeguards. The size target is met; reliable general decision intelligence is not.
Model tree for VynoDePal/laya-mini-cpu-research
Base model
google/bert_uncased_L-6_H-512_A-8