File size: 7,939 Bytes
fbf7b49
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c7d9ed
fbf7b49
 
 
7c7d9ed
 
 
 
 
fbf7b49
 
 
 
7c7d9ed
fbf7b49
 
 
 
 
 
 
 
 
7c7d9ed
 
 
 
fbf7b49
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c7d9ed
 
 
 
 
 
 
 
 
 
 
 
 
fbf7b49
 
 
7c7d9ed
 
 
 
fbf7b49
 
 
7c7d9ed
 
 
 
fbf7b49
 
 
 
 
 
7c7d9ed
 
 
 
 
 
 
 
 
fbf7b49
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c7d9ed
fbf7b49
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c7d9ed
 
 
 
 
fbf7b49
 
 
 
7c7d9ed
fbf7b49
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
---

language:
- ba
license: apache-2.0
pretty_name: BashkirRoBERTa
library_name: transformers
pipeline_tag: fill-mask
tags:
- bashkir
- masked-language-modeling
- roberta
- sentencepiece
- custom-code
- onnx
- onnxruntime
---


# BashkirRoBERTa

> A Bashkir masked language model for fill-mask suggestions and further fine-tuning.

## Overview

BashkirRoBERTa predicts masked SentencePiece tokens from context. It can provide
fill-mask suggestions and serve as an encoder for further fine-tuning, including
the related person-name model. The release includes Transformers weights and FP32,
FP16 and INT8 ONNX graphs. Its custom Pre-LayerNorm architecture requires the
included model code when loaded with Transformers.

| At a glance | |
| --- | --- |
| Task | Masked language modelling / fill-mask |
| Default artifact | `model.safetensors` (Transformers); `onnx/model_int8.onnx` (CPU ONNX) |
| Source | A monolingual Bashkir-language dataset |
| Version / license | v1 / Apache-2.0 |

## Contents

### Files and Configurations

| File | Purpose | Size |
| --- | --- | ---: |
| **`model.safetensors`** | Transformers weights and fine-tuning starting point | 200.2 MB |
| `onnx/model_fp32.onnx` | Full-precision ONNX reference for CPU inference | 242.2 MB |
| `onnx/model_fp16.onnx` | FP16 ONNX model for GPU / DirectML | 121.2 MB |
| **`onnx/model_int8.onnx`** | Compact INT8 ONNX graph for CPU use | 60.9 MB |

| `spm_bashkir_bert_16k.model` | SentencePiece tokenizer | — |

| `config.json` | Model configuration (`auto_map` for custom code) | — |

| `configuration_bashkir_roberta.py`, `modeling_bashkir_roberta.py`, `tokenization_bashkir_roberta.py` | Custom Pre-LayerNorm implementation | — |

| `tokenizer_config.json` | Tokenizer configuration | — |

| `META.json` | Release passport and artifact hashes | — |

| `LICENSE` | Full license text | — |

| `SHA256SUMS` | Release checksums | — |



### Model Architecture



| Property | Value |

| --- | --- |

| Task | Masked language modelling / fill-mask |

| Architecture | Pre-LayerNorm Transformer encoder |

| Transformer blocks | 8 |

| Hidden size / attention heads | 640 / 10 |

| Feed-forward size | 2,560 |

| Context window | 256 subword tokens |

| Parameters | 50.04M |

| Tokenizer | SentencePiece BPE, 16,384 tokens |



The output embedding matrix is tied to the input word embeddings. Token IDs are

fixed: `<pad>` 0, `<unk>` 1, `<s>` 2, `</s>` 3, `[CLS]` 4, `[SEP]` 5 and

`[MASK]` 6.



### Examples



Outputs from the INT8 ONNX model on CPU:



| Input | Top prediction |

| --- | --- |

| `Мин башҡорт телен [MASK].` | `яратам` |

| `Башҡортостан — беҙҙең [MASK].` | `республика` |

| `Өфө — ҙур [MASK].` | `ҡала` |

| `Бөгөн Өфөлә яңы [MASK] асылды.` | `мәсет` |



## Method



The model was pretrained with dynamic masked-language modelling on a monolingual

Bashkir-language dataset assembled from encyclopedic, periodical and literary

sources. The source texts are not distributed in this repository.



```text

raw Bashkir text (monolingual corpus)

  ├── 1. SentencePiece unigram tokenization (vocab size: 32,000)

  ├── 2. dynamic masked-language pretraining (15% subword masking rate)

  ├── 3. export PyTorch Transformers checkpoint

  └── 4. export optimized ONNX graphs: FP32, FP16 & dynamic INT8

```



### Benchmark



On a held-out Bashkir encyclopedic evaluation set the project reports **24.7%
top-1** and **54.0% top-5** accuracy for masked subword prediction. These are
diagnostic MLM results, not a general-purpose language-understanding score: a mask
may represent a whole word or a SentencePiece subword fragment.
The FP32 ONNX graph was checked against the source PyTorch model on multiple
sequence lengths and batch sizes, with matching top-token predictions.

## Quality and Use

Use the Transformers checkpoint for fill-mask experiments or fine-tuning and the
INT8 ONNX graph for compact CPU inference. Choose FP32 ONNX when a full-precision
ONNX reference is needed. Predictions are token suggestions that need review in
context, especially when the mask represents only part of a word.

### Limitations

- Diagnostic MLM accuracy only; not fine-tuned for any downstream task.
- A mask may correspond to a partial subword, not always a full word.
- Predictions reflect the training corpus and may prefer frequent or encyclopedic phrasing.
- No training texts are redistributed; provenance or removal requests go through the maintainer.

## Related Resources

- [BashkirRoBERTa NER](https://huggingface.co/failed09/bashkir-roberta-ner)
  fine-tunes this encoder to find person names and return their text spans. Use it
  when the task is person-name extraction rather than masked-token prediction.

## Usage

```bash

pip install transformers torch huggingface_hub

```

PyTorch (Transformers), which requires `trust_remote_code=True` because of the
custom Pre-LayerNorm architecture:

```python

from transformers import AutoModelForMaskedLM, AutoTokenizer



repo_id = "failed09/bashkir-roberta"

tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)

model = AutoModelForMaskedLM.from_pretrained(repo_id, trust_remote_code=True)



inputs = tokenizer("Мин башҡорт телен [MASK].", return_tensors="pt")

logits = model(**inputs).logits

mask_index = inputs["input_ids"][0].tolist().index(tokenizer.mask_token_id)

prediction_id = logits[0, mask_index].argmax().item()

print(tokenizer.decode([prediction_id]))  # яратам

```

ONNX Runtime for CPU and edge deployment:

```python

import numpy as np

import onnxruntime as ort

import sentencepiece as spm

from huggingface_hub import hf_hub_download



model_path = hf_hub_download("failed09/bashkir-roberta", "onnx/model_int8.onnx")

sp_path = hf_hub_download("failed09/bashkir-roberta", "spm_bashkir_bert_16k.model")



session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])

sp = spm.SentencePieceProcessor(model_file=sp_path)



tokens = [2] + sp.encode("Мин башҡорт телен ") + [6] + sp.encode(".") + [3]

mask_idx = tokens.index(6)

logits = session.run(None, {"input_ids": np.array([tokens], dtype=np.int64)})[0][0, mask_idx]

top_tokens = np.argsort(logits)[::-1][:5]

print([sp.decode([int(t)]) for t in top_tokens])  # ['яратам', 'беләм', 'өйрәнә', ...]

```

For full-precision ONNX inference, change the downloaded filename to
`onnx/model_fp32.onnx`; the input and output names are the same.

## License

The model weights, tokenizer and release code are distributed under the
[Apache-2.0 license](https://huggingface.co/failed09/bashkir-roberta/blob/main/LICENSE). Training texts are not redistributed; their
rights remain with their respective owners. For provenance or removal requests,
contact the maintainer through the Hub.

## Citation

```bibtex

@software{failed09_bashkir_roberta_2026,

  title = {BashkirRoBERTa},

  author = {failed09},

  year = {2026},

  publisher = {Hugging Face},

  url = {https://huggingface.co/failed09/bashkir-roberta},

  note = {Masked language model for Bashkir}

}

```

## Open Bashkir Data and Sources 🐝

This release is part of an open-source effort to support the development,
preservation and practical use of the Bashkir language. Other related models,
datasets and tools are available on the author's Hugging Face profile.

The author does not claim ownership or authorship of the source texts or other
materials used to derive this release; rights and licensing remain with the
original authors, publishers and dataset providers. Source texts are not
redistributed in this repository, so users should follow the licenses and
attribution requirements of the relevant upstream resources.