Bayon (បាយ័ន)
Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km).
The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target.
Bayon serves as the base pretrained model for Bayon Instruct.
Model Summary
| Property | Value |
|---|---|
| Model | Bayon |
| Parameters | ~100M |
| Architecture | Decoder-only Transformer |
| Architecture family | Gemma-style |
| Language | Khmer (km) |
| Tokenizer | Custom Khmer BPE |
| Vocabulary size | 5,000 |
| Pretraining tokens | 1.361B |
| Maximum context length | 2,048 tokens |
| License | Apache-2.0 |
Intended Use
Bayon is intended primarily for:
- Khmer text completion
- Khmer language modeling research
- Research into language-specific model and tokenizer design
- Building Khmer instruction-tuned models
- Studying efficient language modeling for low-resource languages
Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions or instructions.
Training Data
Bayon was pretrained on approximately 1.361 billion Khmer tokens, using the attentionlab/fineweb-2-khmer-extended dataset.
The corpus is based on the Khmer extension of FineWeb-2 and was processed using Khmer-focused data-cleaning and filtering procedures.
The custom tokenizer was used when measuring the reported 1.361B-token training corpus.
For additional information about the training corpus, see the dataset repository:
attentionlab/fineweb-2-khmer-extended
Tokenizer
Bayon uses a custom 5,000-token BPE vocabulary designed specifically for Khmer text.
The tokenizer is intended to provide more compact representations of Khmer text than general-purpose tokenizers that were not specifically optimized for the language.
The tokenizer retains byte fallback. Byte fallback was not observed in the evaluation reported in the associated research.
Because tokenizer efficiency depends on the comparison corpus and tokenizer being evaluated, token-count reductions should not be interpreted as a universal speedup across all workloads.
Architecture
Bayon uses a deep, narrow decoder-only Transformer architecture.
Architecture Configuration
| Hyperparameter | Value |
|---|---|
| Parameters | ~100M |
| Hidden layers | 28 |
| Hidden size | 512 |
| Intermediate size | 2,048 |
| Attention heads | 8 |
| Key/value heads | 2 |
| Head dimension | 64 |
| Maximum position embeddings | 2,048 |
| Vocabulary size | 5,000 |
| RoPE theta | 10,000 |
Architectural Components
Grouped Query Attention (GQA) Bayon uses 8 query heads and 2 key/value heads. This reduces the size of the key/value cache relative to using a separate KV head for every query head.
GeGLU The feed-forward network uses a gated activation structure with GELU activation.
Rotary Position Embeddings (RoPE) Rotary positional embeddings are used to represent token positions within the model's context window.
RMSNorm RMS normalization is used throughout the Transformer architecture.
Pretraining
Bayon was pretrained from scratch using AdamW.
Reported training configuration:
- Optimizer: AdamW
- Weight decay: 0.1
- Peak learning rate: 3e-4
- Final scheduled learning rate: 3e-5
- Warmup: 2,500 steps
- Learning-rate schedule: linear decay
- Gradient clipping: 1.0
- Batch configuration: 4 × 64 × 1,024 tokens
- Planned training: 50,000 steps
- Actual training: stopped at step 18,000
Training was performed using an RTX 4050 and an H100.
The model was stopped at step 18,000 rather than completing the original 50,000-step training plan.
Limitations
Bayon has several important limitations:
- Small model size. At approximately 100M parameters, Bayon has substantially less capacity than many contemporary multilingual language models.
- Base-model behavior. Bayon is not instruction tuned and may not reliably follow instructions or answer questions directly.
- Limited evaluation. The reported benchmark contains 201 Khmer history and culture questions and does not comprehensively measure Khmer language ability.
- Limited training duration. Pretraining stopped at step 18,000 of the planned 50,000 steps.
- Tokenizer-specific measurements. Token-count comparisons depend on the tokenizer and evaluation corpus.
- Data limitations. Web-scale corpora may contain factual errors, duplicated material, unwanted content, or demographic and cultural biases despite filtering.
- No guarantee of factuality. Generated text may be incorrect, fabricated, repetitive, or incoherent.
The model should not be treated as an authoritative source of historical, cultural, legal, medical, or other factual information.
Ethical and Safety Considerations
Bayon is a research language model and has not undergone comprehensive safety evaluation.
Because it was pretrained on web-derived text, it may reproduce undesirable patterns present in its training data, including bias, stereotypes, misinformation, offensive language, or other harmful content.
Applications using Bayon should implement appropriate output filtering and human review where necessary.
Usage
Bayon can be loaded using the Hugging Face Transformers library:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "attentionlab/bayon"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
prompt = "សួស្តី តើអ្នកអាច"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=50,
do_sample=True,
temperature=0.2,
top_k=20,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Citation
Coming Soon
Acknowledgements
Bayon was developed by Seng Sovannpanha as research into language-specific model and tokenizer design for Khmer NLP.
The project aims to contribute openly available resources for research and development in Cambodian/Khmer natural language processing.
- Downloads last month
- 193