Bayon Instruct
Bayon Instruct is an instruction-tuned version of Bayon, a 100M-parameter decoder-only language model designed specifically for Khmer.
Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples.
The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important.
Model Details
| Property | Value |
|---|---|
| Base model | attentionlab/bayon |
| Parameters | ~100M |
| Architecture | Decoder-only Transformer |
| Language | Khmer (km) |
| Tokenizer | Custom 5,000-token Khmer BPE |
| Pretraining data | 1.361B Khmer tokens from FineWeb-2 |
| Instruction-tuning data | 18,000 SFT examples |
| Fine-tuning method | LoRA |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.1 |
| Context/training sequence length | 1,024 tokens |
| License | Apache-2.0 |
The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research.
Intended Use
Bayon Instruct is intended for:
- Khmer instruction following
- Khmer question answering
- Khmer text generation
- Khmer-focused conversational applications
- Research on language-specific and low-resource language models
The model is particularly useful for studying whether a relatively small, language-specific model can maintain Khmer output without relying on a large multilingual model.
Evaluation
On a 201-question Khmer history and culture benchmark, Bayon Instruct achieved:
- 15/201 correct answers
For comparison, Qwen2.5-0.5B-Instruct achieved 4/201 under the reported evaluation setup.
The model also exhibited a low Language Bleed Ratio (LBR) in the reported evaluation:
- Bayon Instruct: 0.06% (13/21,031 tokens)
- SEA-LION 27B: 0.54%
- GPT-OSS 120B: 18.25%
LBR measures the proportion of generated tokens containing characters outside the Khmer Unicode ranges after removing reasoning traces. Because LBR operates on each model's own tokenizer, comparisons across models should be interpreted cautiously: tokenizer vocabulary and tokenization strategy can substantially affect the metric.
The associated research also reports that the custom tokenizer produces substantially fewer tokens than several general-purpose tokenizers on the evaluated Khmer passages.
Language Purity
A primary design goal of Bayon Instruct is Khmer-language consistency.
In the reported three-pass evaluation, only 13 of 21,031 generated tokens were classified as containing non-Khmer characters, corresponding to an LBR of 0.06%.
This result should not be interpreted as evidence that the model is universally more capable or fluent than larger multilingual models. In particular, a Khmer-focused vocabulary inherently constrains which characters and scripts can be generated. Consequently, the observed LBR reflects both the model's learned behavior and the properties of its tokenizer.
Training
Pretraining
Bayon was pretrained from scratch on approximately 1.361 billion Khmer tokens from FineWeb-2.
The reported training configuration included:
- Optimizer: AdamW
- Weight decay: 0.1
- Peak learning rate: 3e-4
- Final learning rate: 3e-5
- Warmup: 2,500 steps
- Schedule: linear decay
- Gradient clipping: 1.0
- Batch configuration: 4 × 64 × 1,024 tokens
- Training stopped at step 18,000 of the planned 50,000 steps
Training was performed across an RTX 4050 and an H100.
Instruction tuning
Bayon was subsequently fine-tuned using LoRA on 18,000 instruction-tuning examples.
Configuration:
- LoRA rank: 16
- LoRA alpha: 32
- LoRA dropout: 0.1
- Target modules: projection layers, embeddings, and language-model head
- Optimizer: AdamW
- Weight decay: 0
- Learning rate: 3e-4
- Warmup ratio: 0.03
- Batch configuration: 8 × 4 × 1,024 tokens
- Epochs: 4
The SFT data was generated using Gemini and was designed to encourage Khmer-focused responses.
Dataset Information
The instruction-tuning dataset is available as attentionlab/bayon-sft.
It contains 18,000 SFT examples generated for instruction tuning Bayon. The data was constructed using Gemini-generated/distilled examples with an emphasis on Khmer-language responses.
Because the same model family and data-generation process are part of the research contribution, users should consider potential teacher-model bias and leakage when interpreting downstream benchmark results.
Limitations
Bayon Instruct is a small, specialized language model and should not be expected to match the general capabilities of substantially larger multilingual models.
Important limitations include:
- Limited evaluation scale. The main factual evaluation contains 201 Khmer history and culture questions.
- Language-purity metrics are not quality metrics. A low LBR indicates low measured language bleeding, but does not establish fluency, factuality, or overall response quality.
- Tokenizer confounding. The Khmer-specific tokenizer restricts the model's output vocabulary and therefore affects LBR directly.
- Limited capability evaluation. The current evaluation does not establish how instruction tuning affects capabilities outside the evaluated Khmer tasks.
- Training-data provenance. The SFT examples were generated using Gemini, so teacher-model biases and possible contamination should be considered.
- Small model size. At approximately 100M parameters, Bayon has substantially less capacity than many modern multilingual instruction-tuned models.
- Benchmark limitations. The 201-question benchmark focuses on Khmer history and culture and should not be treated as a comprehensive measure of Khmer language ability.
Ethical and Safety Considerations
Bayon Instruct is a research model and has not been comprehensively evaluated for harmful, biased, or unsafe outputs.
As with other generative language models, it may produce:
- incorrect or fabricated information;
- culturally inappropriate responses;
- biased or stereotyped content;
- unsafe instructions;
- repetitive or incoherent text.
Outputs should therefore be reviewed before being used in consequential applications.
The model should not be treated as an authoritative source of historical, cultural, legal, medical, or other factual information.
Citation
Coming Soon
Acknowledgements
Bayon was developed as research into native, language-specific model design for Khmer, with an emphasis on understanding the trade-offs between model scale, tokenizer design, and language-specific instruction tuning.
- Downloads last month
- 175
Model tree for attentionlab/bayon-it
Base model
attentionlab/bayon