Slayer149 / README.md
kacperwikiel's picture
Name the model Slayer149
7857dd2 verified
|
Raw History Blame Contribute Delete
6.2 kB
metadata
license: apache-2.0
language:
  - en
library_name: pytorch
pipeline_tag: text-generation
base_model: SlayerLab/gollem-v5-ckpts
tags:
  - gollem
  - tiny-lm
  - muon
  - value-residual
  - continued-pretraining
  - glint-tiny-ml-leaderboard

Slayer149 — 149M English language model

A 148,910,738-parameter English causal language model, grown from the GoLLeM-v5 122.8M final checkpoint and continued for 20,000,014,336 additional tokens. This is a base completion model, not an instruction-tuned assistant.

Evaluation

Full final-checkpoint evaluation, with no best-checkpoint selection:

Metric Result Coverage
BLiMP accuracy 79.47% 67,000 pairs / 67 tasks
ARC-Easy accuracy 54.97% 2,376 test examples
WikiText-2 byte perplexity 2.212573 Full test text
WikiText-2 token perplexity 21.452430 Same evaluation
GLINT overall score, fixed board snapshot 77.11235 Calculated
GLINT efficiency score, fixed board snapshot 77.13585 Calculated

The likelihood protocol follows Glint-Research/Glint-1.3/benchmark.py at c3f99e246aa2c64382f9668dc864533617482d0a: first-256-token truncation for BLiMP and ARC, summed unnormalized log probabilities, zero-shot ARC scoring, and non-overlapping 256-token WikiText contexts. Inference uses BF16 autocast with FP32 log probabilities. Byte perplexity is converted from token perplexity using the measured token/UTF-8-byte ratio. Dataset revisions are pinned in evaluation.json. Loader order and complete task counts are checked.

Against the GLINT board at revision 2c5ea9682babfe2c410ebbbf78113cc98f08a2f9, this would place 4th in overall score and 8th in size-adjusted efficiency among the existing entries plus this model. This is a calculated, self-reported placement, not an accepted leaderboard entry. Competitor scores were not independently re-evaluated here; some competitors use lm-eval-harness and different context/scoring conventions. The ranking and normalized score can change as the board changes. This release does not establish a leaderboard win.

Use

This release uses a small custom PyTorch architecture. Use the included loader; it is not packaged for transformers.AutoModelForCausalLM.

pip install huggingface-hub
hf download SlayerLab/Slayer149 --local-dir Slayer149
cd Slayer149
pip install -r requirements.txt
python generate.py --prompt "The scientific method is" --max-new-tokens 64

A matching CUDA-enabled PyTorch installation is needed for GPU inference/evaluation.

from load_model import load_model
model, tokenizer = load_model('.', device='cuda')
# model(input_ids) returns (logits, optional_loss)

Weights remain FP32 in model.safetensors; tied embeddings are restored by safetensors.torch.load_model. Export verification checks every model tensor and compares logits against the original checkpoint at lengths 16, 256 and 1024. The generation CLI uses BF16 autocast on CUDA, FP32 on CPU, and greedy decoding by default. It recomputes context rather than using a KV cache.

Architecture and training

  • 19 transformer layers; hidden size 768; 12 attention heads; FFN size 2160.
  • Vocabulary 12,288; tied input/output embeddings; context length 1024.
  • RMSNorm, RoPE (theta 100,000), QK normalization, SwiGLU and value residuals.
  • Source: SlayerLab/gollem-v5-ckpts, revision 963440ce6ab4ada7e95da4c1faa28ebb00c082d4, file final_128m_16x768/ckpt_760k.pt (24,903,680,000 source training tokens).
  • FFNs were widened with zero new down-projection columns; three residual blocks were appended with zero output projections. Initial FP32 logits exactly matched the donor on the tested contexts. Optimizer state was reset for continuation.
  • Continuation: 20,000,014,336 tokens, 39,147 actual optimizer updates, seed 1337. Inherited token lineage totals 44,903,694,336; newly added parameters only saw the 20B-token continuation.
  • Cleaned ARC-MIX pool: 9,391,706,576 tokens, reconstructed from the published SlayerLab corpus/removal list. SHA256: ecfd0a4040a0ece6f07f728f092a1b373b6d253b4adfbe259c5107c1fc16681d. Repeated shuffled passes; training corpus is not redistributed here. This work did not independently audit all training/evaluation overlap.
  • Muon for hidden matrices and AdamW for auxiliary parameters; cosine learning rate 2e-4 to 2e-5, 250 reference-unit warmup; Muon LR ratio 33.3333.
  • Eight H100 PCIe GPUs. Effective batch started at 256 sequences, then became 512 after 524,288,000 continuation tokens. Learning rate remained indexed by token position. Grouped Muon, BF16 gradient communication and later CUDA graphs improved throughput. Numerical ordering changed; this is not a bitwise replay of the original trainer.
  • Optimized sustained full-segment throughput: 1.045M tokens/sec including startup, validation and saves; later ordinary windows reached about 1.14M. These are measurements on this host, not expected performance on every H100 node.

training_manifest.json records configuration and hashes. Its reference_step uses 262,144 tokens/unit; it is not the actual optimizer-update count after the batch change. training_metrics.jsonl records the training history.

Reproduce evaluation

python evaluate_glint.py --model-dir . --out reproduced-results.json

This evaluates all three datasets on CUDA using pinned dataset revisions. It can take tens of minutes. board_snapshot.json pins the normalization constants. See release_verification.json for export and official-function parity checks.

Limitations and license

English base model intended for research. It can generate incorrect, repetitive, biased or inappropriate text and is not optimized for instruction following. Small changes in precision, context length, tokenizer handling or benchmark harness can change scores. Overall score and size-adjusted efficiency are distinct.

Weights and GoLLeM implementation: Apache-2.0, following the source model's published license. See LICENSE and NOTICE for source and evaluation attribution.