Bayon Instruct

Bayon Instruct is an instruction-tuned version of Bayon, a 100M-parameter decoder-only language model designed specifically for Khmer.

Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples.

The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important.

Model Details

Property Value
Base model attentionlab/bayon
Parameters ~100M
Architecture Decoder-only Transformer
Language Khmer (km)
Tokenizer Custom 5,000-token Khmer BPE
Pretraining data 1.361B Khmer tokens from FineWeb-2
Instruction-tuning data 18,000 SFT examples
Fine-tuning method LoRA
LoRA rank 16
LoRA alpha 32
LoRA dropout 0.1
Context/training sequence length 1,024 tokens
License Apache-2.0

The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research.

Intended Use

Bayon Instruct is intended for:

  • Khmer instruction following
  • Khmer question answering
  • Khmer text generation
  • Khmer-focused conversational applications
  • Research on language-specific and low-resource language models

The model is particularly useful for studying whether a relatively small, language-specific model can maintain Khmer output without relying on a large multilingual model.

Evaluation

On a 201-question Khmer history and culture benchmark, Bayon Instruct achieved:

  • 15/201 correct answers

For comparison, Qwen2.5-0.5B-Instruct achieved 4/201 under the reported evaluation setup.

The model also exhibited a low Language Bleed Ratio (LBR) in the reported evaluation:

  • Bayon Instruct: 0.06% (13/21,031 tokens)
  • SEA-LION 27B: 0.54%
  • GPT-OSS 120B: 18.25%

LBR measures the proportion of generated tokens containing characters outside the Khmer Unicode ranges after removing reasoning traces. Because LBR operates on each model's own tokenizer, comparisons across models should be interpreted cautiously: tokenizer vocabulary and tokenization strategy can substantially affect the metric.

The associated research also reports that the custom tokenizer produces substantially fewer tokens than several general-purpose tokenizers on the evaluated Khmer passages.

Language Purity

A primary design goal of Bayon Instruct is Khmer-language consistency.

In the reported three-pass evaluation, only 13 of 21,031 generated tokens were classified as containing non-Khmer characters, corresponding to an LBR of 0.06%.

This result should not be interpreted as evidence that the model is universally more capable or fluent than larger multilingual models. In particular, a Khmer-focused vocabulary inherently constrains which characters and scripts can be generated. Consequently, the observed LBR reflects both the model's learned behavior and the properties of its tokenizer.

Training

Pretraining

Bayon was pretrained from scratch on approximately 1.361 billion Khmer tokens from FineWeb-2.

The reported training configuration included:

  • Optimizer: AdamW
  • Weight decay: 0.1
  • Peak learning rate: 3e-4
  • Final learning rate: 3e-5
  • Warmup: 2,500 steps
  • Schedule: linear decay
  • Gradient clipping: 1.0
  • Batch configuration: 4 × 64 × 1,024 tokens
  • Training stopped at step 18,000 of the planned 50,000 steps

Training was performed across an RTX 4050 and an H100.

Instruction tuning

Bayon was subsequently fine-tuned using LoRA on 18,000 instruction-tuning examples.

Configuration:

  • LoRA rank: 16
  • LoRA alpha: 32
  • LoRA dropout: 0.1
  • Target modules: projection layers, embeddings, and language-model head
  • Optimizer: AdamW
  • Weight decay: 0
  • Learning rate: 3e-4
  • Warmup ratio: 0.03
  • Batch configuration: 8 × 4 × 1,024 tokens
  • Epochs: 4

The SFT data was generated using Gemini and was designed to encourage Khmer-focused responses.

Dataset Information

The instruction-tuning dataset is available as attentionlab/bayon-sft.

It contains 18,000 SFT examples generated for instruction tuning Bayon. The data was constructed using Gemini-generated/distilled examples with an emphasis on Khmer-language responses.

Because the same model family and data-generation process are part of the research contribution, users should consider potential teacher-model bias and leakage when interpreting downstream benchmark results.

Limitations

Bayon Instruct is a small, specialized language model and should not be expected to match the general capabilities of substantially larger multilingual models.

Important limitations include:

  1. Limited evaluation scale. The main factual evaluation contains 201 Khmer history and culture questions.
  2. Language-purity metrics are not quality metrics. A low LBR indicates low measured language bleeding, but does not establish fluency, factuality, or overall response quality.
  3. Tokenizer confounding. The Khmer-specific tokenizer restricts the model's output vocabulary and therefore affects LBR directly.
  4. Limited capability evaluation. The current evaluation does not establish how instruction tuning affects capabilities outside the evaluated Khmer tasks.
  5. Training-data provenance. The SFT examples were generated using Gemini, so teacher-model biases and possible contamination should be considered.
  6. Small model size. At approximately 100M parameters, Bayon has substantially less capacity than many modern multilingual instruction-tuned models.
  7. Benchmark limitations. The 201-question benchmark focuses on Khmer history and culture and should not be treated as a comprehensive measure of Khmer language ability.

Ethical and Safety Considerations

Bayon Instruct is a research model and has not been comprehensively evaluated for harmful, biased, or unsafe outputs.

As with other generative language models, it may produce:

  • incorrect or fabricated information;
  • culturally inappropriate responses;
  • biased or stereotyped content;
  • unsafe instructions;
  • repetitive or incoherent text.

Outputs should therefore be reviewed before being used in consequential applications.

The model should not be treated as an authoritative source of historical, cultural, legal, medical, or other factual information.

Citation

Coming Soon

Acknowledgements

Bayon was developed as research into native, language-specific model design for Khmer, with an emphasis on understanding the trade-offs between model scale, tokenizer design, and language-specific instruction tuning.

Downloads last month
175
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for attentionlab/bayon-it

Finetuned
(1)
this model

Dataset used to train attentionlab/bayon-it