Bayon (បាយ័ន)

Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km).

The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target.

Bayon serves as the base pretrained model for Bayon Instruct.

Model Summary

Property Value
Model Bayon
Parameters ~100M
Architecture Decoder-only Transformer
Architecture family Gemma-style
Language Khmer (km)
Tokenizer Custom Khmer BPE
Vocabulary size 5,000
Pretraining tokens 1.361B
Maximum context length 2,048 tokens
License Apache-2.0

Intended Use

Bayon is intended primarily for:

  • Khmer text completion
  • Khmer language modeling research
  • Research into language-specific model and tokenizer design
  • Building Khmer instruction-tuned models
  • Studying efficient language modeling for low-resource languages

Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions or instructions.

Training Data

Bayon was pretrained on approximately 1.361 billion Khmer tokens, using the attentionlab/fineweb-2-khmer-extended dataset.

The corpus is based on the Khmer extension of FineWeb-2 and was processed using Khmer-focused data-cleaning and filtering procedures.

The custom tokenizer was used when measuring the reported 1.361B-token training corpus.

For additional information about the training corpus, see the dataset repository:

attentionlab/fineweb-2-khmer-extended

Tokenizer

Bayon uses a custom 5,000-token BPE vocabulary designed specifically for Khmer text.

The tokenizer is intended to provide more compact representations of Khmer text than general-purpose tokenizers that were not specifically optimized for the language.

The tokenizer retains byte fallback. Byte fallback was not observed in the evaluation reported in the associated research.

Because tokenizer efficiency depends on the comparison corpus and tokenizer being evaluated, token-count reductions should not be interpreted as a universal speedup across all workloads.

Architecture

Bayon uses a deep, narrow decoder-only Transformer architecture.

Architecture Configuration

Hyperparameter Value
Parameters ~100M
Hidden layers 28
Hidden size 512
Intermediate size 2,048
Attention heads 8
Key/value heads 2
Head dimension 64
Maximum position embeddings 2,048
Vocabulary size 5,000
RoPE theta 10,000

Architectural Components

Grouped Query Attention (GQA) Bayon uses 8 query heads and 2 key/value heads. This reduces the size of the key/value cache relative to using a separate KV head for every query head.

GeGLU The feed-forward network uses a gated activation structure with GELU activation.

Rotary Position Embeddings (RoPE) Rotary positional embeddings are used to represent token positions within the model's context window.

RMSNorm RMS normalization is used throughout the Transformer architecture.

Pretraining

Bayon was pretrained from scratch using AdamW.

Reported training configuration:

  • Optimizer: AdamW
  • Weight decay: 0.1
  • Peak learning rate: 3e-4
  • Final scheduled learning rate: 3e-5
  • Warmup: 2,500 steps
  • Learning-rate schedule: linear decay
  • Gradient clipping: 1.0
  • Batch configuration: 4 × 64 × 1,024 tokens
  • Planned training: 50,000 steps
  • Actual training: stopped at step 18,000

Training was performed using an RTX 4050 and an H100.

The model was stopped at step 18,000 rather than completing the original 50,000-step training plan.

Limitations

Bayon has several important limitations:

  1. Small model size. At approximately 100M parameters, Bayon has substantially less capacity than many contemporary multilingual language models.
  2. Base-model behavior. Bayon is not instruction tuned and may not reliably follow instructions or answer questions directly.
  3. Limited evaluation. The reported benchmark contains 201 Khmer history and culture questions and does not comprehensively measure Khmer language ability.
  4. Limited training duration. Pretraining stopped at step 18,000 of the planned 50,000 steps.
  5. Tokenizer-specific measurements. Token-count comparisons depend on the tokenizer and evaluation corpus.
  6. Data limitations. Web-scale corpora may contain factual errors, duplicated material, unwanted content, or demographic and cultural biases despite filtering.
  7. No guarantee of factuality. Generated text may be incorrect, fabricated, repetitive, or incoherent.

The model should not be treated as an authoritative source of historical, cultural, legal, medical, or other factual information.

Ethical and Safety Considerations

Bayon is a research language model and has not undergone comprehensive safety evaluation.

Because it was pretrained on web-derived text, it may reproduce undesirable patterns present in its training data, including bias, stereotypes, misinformation, offensive language, or other harmful content.

Applications using Bayon should implement appropriate output filtering and human review where necessary.

Usage

Bayon can be loaded using the Hugging Face Transformers library:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "attentionlab/bayon"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

prompt = "សួស្តី តើអ្នកអាច"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=50,
        do_sample=True,
        temperature=0.2,
        top_k=20,
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Citation

Coming Soon

Acknowledgements

Bayon was developed by Seng Sovannpanha as research into language-specific model and tokenizer design for Khmer NLP.

The project aims to contribute openly available resources for research and development in Cambodian/Khmer natural language processing.

Downloads last month
193
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for attentionlab/bayon

Finetunes
1 model
Quantizations
1 model

Dataset used to train attentionlab/bayon