--- license: apache-2.0 language: - en library_name: pytorch pipeline_tag: text-generation base_model: SlayerLab/gollem-v5-ckpts tags: - gollem - tiny-lm - muon - value-residual - continued-pretraining - glint-tiny-ml-leaderboard --- # Slayer149 — 149M English language model A **148,910,738-parameter English causal language model**, grown from the GoLLeM-v5 122.8M final checkpoint and continued for **20,000,014,336 additional tokens**. This is a base completion model, not an instruction-tuned assistant. ## Evaluation Full final-checkpoint evaluation, with no best-checkpoint selection: | Metric | Result | Coverage | |---|---:|---| | BLiMP accuracy | **79.47%** | 67,000 pairs / 67 tasks | | ARC-Easy accuracy | **54.97%** | 2,376 test examples | | WikiText-2 byte perplexity | **2.212573** | Full test text | | WikiText-2 token perplexity | 21.452430 | Same evaluation | | GLINT overall score, fixed board snapshot | **77.11235** | Calculated | | GLINT efficiency score, fixed board snapshot | **77.13585** | Calculated | The likelihood protocol follows `Glint-Research/Glint-1.3/benchmark.py` at `c3f99e246aa2c64382f9668dc864533617482d0a`: first-256-token truncation for BLiMP and ARC, summed unnormalized log probabilities, zero-shot ARC scoring, and non-overlapping 256-token WikiText contexts. Inference uses BF16 autocast with FP32 log probabilities. Byte perplexity is converted from token perplexity using the measured token/UTF-8-byte ratio. Dataset revisions are pinned in [evaluation.json](evaluation.json). Loader order and complete task counts are checked. Against the [GLINT board](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard) at revision `2c5ea9682babfe2c410ebbbf78113cc98f08a2f9`, this would place **4th in overall score** and **8th in size-adjusted efficiency** among the existing entries plus this model. This is a calculated, self-reported placement, not an accepted leaderboard entry. Competitor scores were not independently re-evaluated here; some competitors use lm-eval-harness and different context/scoring conventions. The ranking and normalized score can change as the board changes. This release does not establish a leaderboard win. ## Use This release uses a small custom PyTorch architecture. Use the included loader; it is not packaged for `transformers.AutoModelForCausalLM`. ```bash pip install huggingface-hub hf download SlayerLab/Slayer149 --local-dir Slayer149 cd Slayer149 pip install -r requirements.txt python generate.py --prompt "The scientific method is" --max-new-tokens 64 ``` A matching CUDA-enabled PyTorch installation is needed for GPU inference/evaluation. ```python from load_model import load_model model, tokenizer = load_model('.', device='cuda') # model(input_ids) returns (logits, optional_loss) ``` Weights remain FP32 in `model.safetensors`; tied embeddings are restored by `safetensors.torch.load_model`. Export verification checks every model tensor and compares logits against the original checkpoint at lengths 16, 256 and 1024. The generation CLI uses BF16 autocast on CUDA, FP32 on CPU, and greedy decoding by default. It recomputes context rather than using a KV cache. ## Architecture and training - 19 transformer layers; hidden size 768; 12 attention heads; FFN size 2160. - Vocabulary 12,288; tied input/output embeddings; context length 1024. - RMSNorm, RoPE (theta 100,000), QK normalization, SwiGLU and value residuals. - Source: `SlayerLab/gollem-v5-ckpts`, revision `963440ce6ab4ada7e95da4c1faa28ebb00c082d4`, file `final_128m_16x768/ckpt_760k.pt` (24,903,680,000 source training tokens). - FFNs were widened with zero new down-projection columns; three residual blocks were appended with zero output projections. Initial FP32 logits exactly matched the donor on the tested contexts. Optimizer state was reset for continuation. - Continuation: 20,000,014,336 tokens, 39,147 actual optimizer updates, seed 1337. Inherited token lineage totals 44,903,694,336; newly added parameters only saw the 20B-token continuation. - Cleaned ARC-MIX pool: 9,391,706,576 tokens, reconstructed from the published SlayerLab corpus/removal list. SHA256: `ecfd0a4040a0ece6f07f728f092a1b373b6d253b4adfbe259c5107c1fc16681d`. Repeated shuffled passes; training corpus is not redistributed here. This work did not independently audit all training/evaluation overlap. - Muon for hidden matrices and AdamW for auxiliary parameters; cosine learning rate 2e-4 to 2e-5, 250 reference-unit warmup; Muon LR ratio 33.3333. - Eight H100 PCIe GPUs. Effective batch started at 256 sequences, then became 512 after 524,288,000 continuation tokens. Learning rate remained indexed by token position. Grouped Muon, BF16 gradient communication and later CUDA graphs improved throughput. Numerical ordering changed; this is not a bitwise replay of the original trainer. - Optimized sustained full-segment throughput: 1.045M tokens/sec including startup, validation and saves; later ordinary windows reached about 1.14M. These are measurements on this host, not expected performance on every H100 node. `training_manifest.json` records configuration and hashes. Its `reference_step` uses 262,144 tokens/unit; it is not the actual optimizer-update count after the batch change. `training_metrics.jsonl` records the training history. ## Reproduce evaluation ```bash python evaluate_glint.py --model-dir . --out reproduced-results.json ``` This evaluates all three datasets on CUDA using pinned dataset revisions. It can take tens of minutes. `board_snapshot.json` pins the normalization constants. See `release_verification.json` for export and official-function parity checks. ## Limitations and license English base model intended for research. It can generate incorrect, repetitive, biased or inappropriate text and is not optimized for instruction following. Small changes in precision, context length, tokenizer handling or benchmark harness can change scores. Overall score and size-adjusted efficiency are distinct. Weights and GoLLeM implementation: Apache-2.0, following the source model's published license. See `LICENSE` and `NOTICE` for source and evaluation attribution.