Llama-1B Clean BPE
Llama-1B Clean BPE is a decoder-only English language model trained from random initialization for controlled tokenizer research. It uses a 32k byte-level BPE tokenizer and is not instruction tuned.
Training
The model was trained on approximately 10 billion tokenizer tokens from FineWeb-EDU-dedup-10B. The training run used a fixed seed (42), a cosine learning-rate schedule, and a maximum sequence length of 2,048 tokens.
Architecture
- Llama-style causal language model with approximately 1B parameters
- 16 transformer layers
- Hidden size 2,048; MLP size 8,192
- 32 attention heads and 8 key/value heads
- 2,048-token context length
- 32,002-token vocabulary, including padding and end-of-sequence tokens
- Tied input and output embeddings; bfloat16 weights
Intended use
This checkpoint is intended for research on language modeling and tokenization. It can be loaded with AutoModelForCausalLM and its bundled tokenizer with AutoTokenizer.
It is a base language model, not a chat model or instruction-following model. It has not been evaluated for safety, factuality, fairness, or production use.
- Downloads last month
- 193