Bashkir Full Fine-Tuning Models

Full fine-tuned GPT-2 family models on Bashkir Web Corpus (71,567 documents).

Part of the BashkirNLP project. See the paper for details.

Structure

Each top-level subfolder <model>_seed<N>/ contains a full fine-tuned language model. The classification/ folder contains sequence-classification heads trained on top of the same base models.

distilgpt2_seed42/
distilgpt2_seed123/
distilgpt2_seed2024/
gpt2_seed42/
...
classification/
    cls_binary_distilgpt2_seed42/
    cls_multiclass_distilgpt2_seed42/
    ...

Training

  • Base models: distilgpt2 (82M), gpt2 (124M), gpt2-medium (355M)
  • Corpus: Bashkir Web Corpus (57,254 train / 7,157 val / 7,156 test)
  • Epochs: 3
  • Seeds: 42, 123, 2024
  • max_seq_len: 128

Evaluation

See step1_report.txt in the paper repository for metrics (PPL, BPB, catastrophic forgetting, target-language ratio).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support