Bashkir Full Fine-Tuning Models
Full fine-tuned GPT-2 family models on Bashkir Web Corpus (71,567 documents).
Part of the BashkirNLP project. See the paper for details.
Structure
Each top-level subfolder <model>_seed<N>/ contains a full fine-tuned language model.
The classification/ folder contains sequence-classification heads trained on top of the same base models.
distilgpt2_seed42/
distilgpt2_seed123/
distilgpt2_seed2024/
gpt2_seed42/
...
classification/
cls_binary_distilgpt2_seed42/
cls_multiclass_distilgpt2_seed42/
...
Training
- Base models:
distilgpt2(82M),gpt2(124M),gpt2-medium(355M) - Corpus: Bashkir Web Corpus (57,254 train / 7,157 val / 7,156 test)
- Epochs: 3
- Seeds: 42, 123, 2024
max_seq_len: 128
Evaluation
See step1_report.txt in the paper repository for metrics (PPL, BPB, catastrophic forgetting, target-language ratio).
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support