Llama-1B Clean BPE

Llama-1B Clean BPE is a decoder-only English language model trained from random initialization for controlled tokenizer research. It uses a 32k byte-level BPE tokenizer and is not instruction tuned.

Training

The model was trained on approximately 10 billion tokenizer tokens from FineWeb-EDU-dedup-10B. The training run used a fixed seed (42), a cosine learning-rate schedule, and a maximum sequence length of 2,048 tokens.

Architecture

  • Llama-style causal language model with approximately 1B parameters
  • 16 transformer layers
  • Hidden size 2,048; MLP size 8,192
  • 32 attention heads and 8 key/value heads
  • 2,048-token context length
  • 32,002-token vocabulary, including padding and end-of-sequence tokens
  • Tied input and output embeddings; bfloat16 weights

Intended use

This checkpoint is intended for research on language modeling and tokenization. It can be loaded with AutoModelForCausalLM and its bundled tokenizer with AutoTokenizer.

It is a base language model, not a chat model or instruction-following model. It has not been evaluated for safety, factuality, fairness, or production use.

Downloads last month
193
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support