EXLLM / DATA_PROVENANCE.md
ToTo-40417
Publish EXQ12 and 5M lineage metadata
b41d6e3 unverified
|
Raw History Blame Contribute Delete
2.56 kB

Data and Weight Provenance

Model lineage

EXLLM-0.005B-Instruct is a project-original model with 0.005377824B unique trainable parameters.

  1. The EXLLM project originates from random initialization.
  2. The current 0.005B release was expanded from the published project-owned weights/EXLLM-v1.0.0.safetensors checkpoint by transplanting dimension-compatible parameters.
  3. Additional staged training used only project-generated Japanese instruction/response data.
  4. Corrective datasets were added in response to regression failures.
  5. The final fp32 checkpoint was exported to EXLLM8 and EXQ12 deployment formats.

No external pretrained checkpoint was used. The model is not a distillation or quantization of a third-party LLM.

The direct parent hash and every 5M training stage are recorded in training/release-5m-stages.json. The source repository contains the complete datasets and tools/replay_5m.py. EXQ12 is deterministically generated from EXLLM8 by tools/export_exq12.py.

Training data

The project contains 0.086408M JSONL records across its published staged files. This number is the sum of file records, not a deduplicated training-example count. The files include overlapping base, validation, robustness, recovery, balance, consistency, and final corrective stages.

The data covers:

  • short Japanese greetings and UI interaction;
  • EXLLM and developer identity responses;
  • short definitions of common terms;
  • offline and current-information limitations;
  • unknown-input fallback behavior;
  • mixed Unicode and UTF-8 byte fallback;
  • deterministic calculator routing;
  • regression cases discovered during development.

No external public text dataset was imported into the published project corpus.

Published checkpoint metadata

Field Value
Parameters 0.005377824B
Global step 1,810
Initialization Published project-owned v1.0 checkpoint transplantation
Final corrective source EXLLM-v1.1-5m-release2.pt
Final corrective data v1_1_release_fix.jsonl
Final corrective steps 100
Final corrective learning rate 6e-6
Final corrective loss 0.03875 → 0.02420

Licensing

The released model weights, project-generated data, and reference code are distributed under Apache License 2.0 unless otherwise stated.

Scope

The corpus is intentionally narrow. Its presence in the training data does not establish comprehensive knowledge, factual authority, or currency. See the model card for intended use and limitations.