failed09 commited on
Commit
fbf7b49
·
verified ·
1 Parent(s): 078b576

Release update

Browse files
Files changed (4) hide show
  1. LICENSE +202 -0
  2. META.json +1 -1
  3. README.md +184 -182
  4. SHA256SUMS +3 -2
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
META.json CHANGED
@@ -4,7 +4,7 @@
4
  "type": "masked_language_model",
5
  "status": "current",
6
  "language": "ba",
7
- "license": "other",
8
  "source": "a monolingual Bashkir-language dataset",
9
  "created_at": "2026-09-17T18:20:06Z",
10
  "architecture": "Pre-LayerNorm Transformer encoder",
 
4
  "type": "masked_language_model",
5
  "status": "current",
6
  "language": "ba",
7
+ "license": "apache-2.0",
8
  "source": "a monolingual Bashkir-language dataset",
9
  "created_at": "2026-09-17T18:20:06Z",
10
  "architecture": "Pre-LayerNorm Transformer encoder",
README.md CHANGED
@@ -1,182 +1,184 @@
1
- ---
2
- language:
3
- - ba
4
- license: other
5
- pretty_name: BashkirRoBERTa
6
- library_name: transformers
7
- pipeline_tag: fill-mask
8
- tags:
9
- - bashkir
10
- - masked-language-modeling
11
- - roberta
12
- - sentencepiece
13
- - custom-code
14
- - onnx
15
- - onnxruntime
16
- ---
17
-
18
- # BashkirRoBERTa
19
-
20
- > A masked language model for Bashkir, for fill-mask, spellchecking and
21
- > foundation fine-tuning.
22
-
23
- ## Overview
24
-
25
- A masked language model for Bashkir. Given a sentence with one `[MASK]` token, it
26
- predicts the most probable missing Bashkir token from context. The model is useful
27
- for fill-mask experiments, spellchecking and as a foundation for further
28
- fine-tuning. It preserves a custom Pre-LayerNorm architecture rather than the
29
- stock post-LayerNorm RoBERTa implementation.
30
-
31
- | At a glance | |
32
- | --- | --- |
33
- | Task | Masked language modelling / fill-mask |
34
- | Default artifact | `model.safetensors` (Transformers) or `onnx/model_int8.onnx` (ONNX) |
35
- | Source | A monolingual Bashkir-language dataset |
36
- | Version / license | v1 / custom terms (`other`) |
37
-
38
- ## Contents
39
-
40
- ### Files and Configurations
41
-
42
- | File | Purpose | Size |
43
- | --- | --- | ---: |
44
- | `model.safetensors` | PyTorch weights for Transformers | 200.2 MB |
45
- | `onnx/model_fp16.onnx` | FP16 ONNX model for GPU / DirectML | 121.2 MB |
46
- | `onnx/model_int8.onnx` | INT8 ONNX model for fast CPU / mobile | 60.9 MB |
47
- | `spm_bashkir_bert_16k.model` | SentencePiece tokenizer | — |
48
- | `config.json` | Model configuration (`auto_map` for custom code) | — |
49
- | `configuration_bashkir_roberta.py`, `modeling_bashkir_roberta.py`, `tokenization_bashkir_roberta.py` | Custom Pre-LayerNorm implementation | — |
50
- | `tokenizer_config.json` | Tokenizer configuration | — |
51
- | `META.json` | Release passport and artifact hashes | — |
52
- | `SHA256SUMS` | Release checksums | — |
53
-
54
- ### Model Architecture
55
-
56
- | Property | Value |
57
- | --- | --- |
58
- | Task | Masked language modelling / fill-mask |
59
- | Architecture | Pre-LayerNorm Transformer encoder |
60
- | Transformer blocks | 8 |
61
- | Hidden size / attention heads | 640 / 10 |
62
- | Feed-forward size | 2,560 |
63
- | Context window | 256 subword tokens |
64
- | Parameters | 50.04M |
65
- | Tokenizer | SentencePiece BPE, 16,384 tokens |
66
-
67
- The output embedding matrix is tied to the input word embeddings. Token IDs are
68
- fixed: `<pad>` 0, `<unk>` 1, `<s>` 2, `</s>` 3, `[CLS]` 4, `[SEP]` 5 and
69
- `[MASK]` 6.
70
-
71
- ### Examples
72
-
73
- Outputs from the INT8 ONNX model on CPU:
74
-
75
- | Input | Top prediction |
76
- | --- | --- |
77
- | `Мин башҡорт телен [MASK].` | `яратам` |
78
- | `Башҡортостан — беҙҙең [MASK].` | `республика` |
79
- | `Өфө — ҙур [MASK].` | `ҡала` |
80
- | `Бөгөн Өфөлә яңы [MASK] асылды.` | `мәсет` |
81
-
82
- ## Method
83
-
84
- The model was pretrained with dynamic masked-language modelling on a monolingual
85
- Bashkir-language dataset assembled from encyclopedic, periodical and literary
86
- sources. The source texts are not distributed in this repository.
87
-
88
- ### Evaluation
89
-
90
- On a held-out Bashkir encyclopedic evaluation set the project reports **24.7%
91
- top-1** and **54.0% top-5** accuracy for masked subword prediction. These are
92
- diagnostic MLM results, not a general-purpose language-understanding score: a mask
93
- may represent a whole word or a SentencePiece subword fragment.
94
-
95
- ## Quality and Use
96
-
97
- This is a research model, not a production language service. Fill-mask predictions
98
- are ranking suggestions that require context-appropriate review, especially for
99
- ambiguous or short contexts. The checkpoint is released under custom terms while
100
- the source-rights audit is completed.
101
-
102
- ### Limitations
103
-
104
- - Diagnostic MLM accuracy only; not fine-tuned for any downstream task.
105
- - A mask may correspond to a partial subword, not always a full word.
106
- - Predictions reflect the training corpus and may prefer frequent or encyclopedic phrasing.
107
- - No training texts are redistributed; provenance or removal requests go through the maintainer.
108
-
109
- ## Usage
110
-
111
- ```bash
112
- pip install transformers torch huggingface_hub
113
- ```
114
-
115
- PyTorch (Transformers), which requires `trust_remote_code=True` because of the
116
- custom Pre-LayerNorm architecture:
117
-
118
- ```python
119
- from transformers import AutoModelForMaskedLM, AutoTokenizer
120
-
121
- repo_id = "failed09/bashkir-roberta"
122
- tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
123
- model = AutoModelForMaskedLM.from_pretrained(repo_id, trust_remote_code=True)
124
-
125
- inputs = tokenizer("Мин башҡорт телен [MASK].", return_tensors="pt")
126
- logits = model(**inputs).logits
127
- mask_index = inputs["input_ids"][0].tolist().index(tokenizer.mask_token_id)
128
- prediction_id = logits[0, mask_index].argmax().item()
129
- print(tokenizer.decode([prediction_id])) # яратам
130
- ```
131
-
132
- ONNX Runtime for CPU and edge deployment:
133
-
134
- ```python
135
- import numpy as np
136
- import onnxruntime as ort
137
- import sentencepiece as spm
138
- from huggingface_hub import hf_hub_download
139
-
140
- model_path = hf_hub_download("failed09/bashkir-roberta", "onnx/model_int8.onnx")
141
- sp_path = hf_hub_download("failed09/bashkir-roberta", "spm_bashkir_bert_16k.model")
142
-
143
- session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
144
- sp = spm.SentencePieceProcessor(model_file=sp_path)
145
-
146
- tokens = [2] + sp.encode("Мин башҡорт телен ") + [6] + sp.encode(".") + [3]
147
- mask_idx = tokens.index(6)
148
- logits = session.run(None, {"input_ids": np.array([tokens], dtype=np.int64)})[0][0, mask_idx]
149
- top_tokens = np.argsort(logits)[::-1][:5]
150
- print([sp.decode([int(t)]) for t in top_tokens]) # ['яратам', 'беләм', 'өйрәнә', ...]
151
- ```
152
-
153
- ## License
154
-
155
- The checkpoint is released under custom terms (`other` on the Hub) while the
156
- source-rights audit is completed. No training texts are redistributed. For
157
- provenance or removal requests, contact the maintainer through the Hub.
158
-
159
- ## Citation
160
-
161
- ```bibtex
162
- @software{failed09_bashkir_roberta_2026,
163
- title = {BashkirRoBERTa},
164
- author = {failed09},
165
- year = {2026},
166
- publisher = {Hugging Face},
167
- url = {https://huggingface.co/failed09/bashkir-roberta},
168
- note = {Masked language model for Bashkir}
169
- }
170
- ```
171
-
172
- ## Open Bashkir Data and Sources 🐝
173
-
174
- This release is part of an open-source effort to support the development,
175
- preservation and practical use of the Bashkir language. Other related models,
176
- datasets and tools are available on the author's Hugging Face profile.
177
-
178
- The author does not claim ownership or authorship of the source texts or other
179
- materials used to derive this release; rights and licensing remain with the
180
- original authors, publishers and dataset providers. Source texts are not
181
- redistributed in this repository, so users should follow the licenses and
182
- attribution requirements of the relevant upstream resources.
 
 
 
1
+ ---
2
+ language:
3
+ - ba
4
+ license: apache-2.0
5
+ pretty_name: BashkirRoBERTa
6
+ library_name: transformers
7
+ pipeline_tag: fill-mask
8
+ tags:
9
+ - bashkir
10
+ - masked-language-modeling
11
+ - roberta
12
+ - sentencepiece
13
+ - custom-code
14
+ - onnx
15
+ - onnxruntime
16
+ ---
17
+
18
+ # BashkirRoBERTa
19
+
20
+ > A masked language model for Bashkir, for fill-mask, spellchecking and
21
+ > foundation fine-tuning.
22
+
23
+ ## Overview
24
+
25
+ A masked language model for Bashkir. Given a sentence with one `[MASK]` token, it
26
+ predicts the most probable missing Bashkir token from context. The model is useful
27
+ for fill-mask experiments, spellchecking and as a foundation for further
28
+ fine-tuning. It preserves a custom Pre-LayerNorm architecture rather than the
29
+ stock post-LayerNorm RoBERTa implementation.
30
+
31
+ | At a glance | |
32
+ | --- | --- |
33
+ | Task | Masked language modelling / fill-mask |
34
+ | Default artifact | `model.safetensors` (Transformers) or `onnx/model_int8.onnx` (ONNX) |
35
+ | Source | A monolingual Bashkir-language dataset |
36
+ | Version / license | v1 / Apache-2.0 |
37
+
38
+ ## Contents
39
+
40
+ ### Files and Configurations
41
+
42
+ | File | Purpose | Size |
43
+ | --- | --- | ---: |
44
+ | `model.safetensors` | PyTorch weights for Transformers | 200.2 MB |
45
+ | `onnx/model_fp16.onnx` | FP16 ONNX model for GPU / DirectML | 121.2 MB |
46
+ | `onnx/model_int8.onnx` | INT8 ONNX model for fast CPU / mobile | 60.9 MB |
47
+ | `spm_bashkir_bert_16k.model` | SentencePiece tokenizer | — |
48
+ | `config.json` | Model configuration (`auto_map` for custom code) | — |
49
+ | `configuration_bashkir_roberta.py`, `modeling_bashkir_roberta.py`, `tokenization_bashkir_roberta.py` | Custom Pre-LayerNorm implementation | — |
50
+ | `tokenizer_config.json` | Tokenizer configuration | — |
51
+ | `META.json` | Release passport and artifact hashes | — |
52
+ | `LICENSE` | Full license text | — |
53
+ | `SHA256SUMS` | Release checksums | — |
54
+
55
+ ### Model Architecture
56
+
57
+ | Property | Value |
58
+ | --- | --- |
59
+ | Task | Masked language modelling / fill-mask |
60
+ | Architecture | Pre-LayerNorm Transformer encoder |
61
+ | Transformer blocks | 8 |
62
+ | Hidden size / attention heads | 640 / 10 |
63
+ | Feed-forward size | 2,560 |
64
+ | Context window | 256 subword tokens |
65
+ | Parameters | 50.04M |
66
+ | Tokenizer | SentencePiece BPE, 16,384 tokens |
67
+
68
+ The output embedding matrix is tied to the input word embeddings. Token IDs are
69
+ fixed: `<pad>` 0, `<unk>` 1, `<s>` 2, `</s>` 3, `[CLS]` 4, `[SEP]` 5 and
70
+ `[MASK]` 6.
71
+
72
+ ### Examples
73
+
74
+ Outputs from the INT8 ONNX model on CPU:
75
+
76
+ | Input | Top prediction |
77
+ | --- | --- |
78
+ | `Мин башҡорт телен [MASK].` | `яратам` |
79
+ | `Башҡортостан — беҙҙең [MASK].` | `республика` |
80
+ | `Өфө — ҙур [MASK].` | `ҡала` |
81
+ | `Бөгөн Өфөлә яңы [MASK] асылды.` | `мәсет` |
82
+
83
+ ## Method
84
+
85
+ The model was pretrained with dynamic masked-language modelling on a monolingual
86
+ Bashkir-language dataset assembled from encyclopedic, periodical and literary
87
+ sources. The source texts are not distributed in this repository.
88
+
89
+ ### Evaluation
90
+
91
+ On a held-out Bashkir encyclopedic evaluation set the project reports **24.7%
92
+ top-1** and **54.0% top-5** accuracy for masked subword prediction. These are
93
+ diagnostic MLM results, not a general-purpose language-understanding score: a mask
94
+ may represent a whole word or a SentencePiece subword fragment.
95
+
96
+ ## Quality and Use
97
+
98
+ This is a research model, not a production language service. Fill-mask predictions
99
+ are ranking suggestions that require context-appropriate review, especially for
100
+ ambiguous or short contexts. The model weights and release code are available
101
+ under Apache-2.0; source-text provenance remains documented separately.
102
+
103
+ ### Limitations
104
+
105
+ - Diagnostic MLM accuracy only; not fine-tuned for any downstream task.
106
+ - A mask may correspond to a partial subword, not always a full word.
107
+ - Predictions reflect the training corpus and may prefer frequent or encyclopedic phrasing.
108
+ - No training texts are redistributed; provenance or removal requests go through the maintainer.
109
+
110
+ ## Usage
111
+
112
+ ```bash
113
+ pip install transformers torch huggingface_hub
114
+ ```
115
+
116
+ PyTorch (Transformers), which requires `trust_remote_code=True` because of the
117
+ custom Pre-LayerNorm architecture:
118
+
119
+ ```python
120
+ from transformers import AutoModelForMaskedLM, AutoTokenizer
121
+
122
+ repo_id = "failed09/bashkir-roberta"
123
+ tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
124
+ model = AutoModelForMaskedLM.from_pretrained(repo_id, trust_remote_code=True)
125
+
126
+ inputs = tokenizer("Мин башҡорт телен [MASK].", return_tensors="pt")
127
+ logits = model(**inputs).logits
128
+ mask_index = inputs["input_ids"][0].tolist().index(tokenizer.mask_token_id)
129
+ prediction_id = logits[0, mask_index].argmax().item()
130
+ print(tokenizer.decode([prediction_id])) # яратам
131
+ ```
132
+
133
+ ONNX Runtime for CPU and edge deployment:
134
+
135
+ ```python
136
+ import numpy as np
137
+ import onnxruntime as ort
138
+ import sentencepiece as spm
139
+ from huggingface_hub import hf_hub_download
140
+
141
+ model_path = hf_hub_download("failed09/bashkir-roberta", "onnx/model_int8.onnx")
142
+ sp_path = hf_hub_download("failed09/bashkir-roberta", "spm_bashkir_bert_16k.model")
143
+
144
+ session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
145
+ sp = spm.SentencePieceProcessor(model_file=sp_path)
146
+
147
+ tokens = [2] + sp.encode("Мин башҡорт телен ") + [6] + sp.encode(".") + [3]
148
+ mask_idx = tokens.index(6)
149
+ logits = session.run(None, {"input_ids": np.array([tokens], dtype=np.int64)})[0][0, mask_idx]
150
+ top_tokens = np.argsort(logits)[::-1][:5]
151
+ print([sp.decode([int(t)]) for t in top_tokens]) # ['яратам', 'беләм', 'өйрәнә', ...]
152
+ ```
153
+
154
+ ## License
155
+
156
+ The model weights, tokenizer and release code are distributed under the
157
+ [Apache-2.0 license](LICENSE). Training texts are not redistributed; their
158
+ rights remain with their respective owners. For provenance or removal requests,
159
+ contact the maintainer through the Hub.
160
+
161
+ ## Citation
162
+
163
+ ```bibtex
164
+ @software{failed09_bashkir_roberta_2026,
165
+ title = {BashkirRoBERTa},
166
+ author = {failed09},
167
+ year = {2026},
168
+ publisher = {Hugging Face},
169
+ url = {https://huggingface.co/failed09/bashkir-roberta},
170
+ note = {Masked language model for Bashkir}
171
+ }
172
+ ```
173
+
174
+ ## Open Bashkir Data and Sources 🐝
175
+
176
+ This release is part of an open-source effort to support the development,
177
+ preservation and practical use of the Bashkir language. Other related models,
178
+ datasets and tools are available on the author's Hugging Face profile.
179
+
180
+ The author does not claim ownership or authorship of the source texts or other
181
+ materials used to derive this release; rights and licensing remain with the
182
+ original authors, publishers and dataset providers. Source texts are not
183
+ redistributed in this repository, so users should follow the licenses and
184
+ attribution requirements of the relevant upstream resources.
SHA256SUMS CHANGED
@@ -1,5 +1,6 @@
1
- a56541821a9afbbacaab9e6ce0322748543b3ada068e42956b8993f927ddea27 META.json
2
- b077d3131f87aa776cecd0c4feb214686c9dc961a0f76cd15658f5f2a6c2fc96 README.md
 
3
  59b773953af5b2a69adb7e3c5d345a90f34f848f6a1543309d5acc7fb78b28c4 config.json
4
  e9f179aad7619153724920c86a90ecc3c161fd15ffc3332f200e9f2073a07f72 configuration_bashkir_roberta.py
5
  48e4d4e7b8af675d222e58e14776abac24b7e83a937d7884bd9d735ea9dd2173 model.safetensors
 
1
+ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
2
+ f6d59748711ddcc45af84d09cf7970a276ee93820a83dda4fb5fda8209f878ca META.json
3
+ e55a82e7ce32715e0d71e3b9fc809a623ad1d987b4c39e7458534d34439ff149 README.md
4
  59b773953af5b2a69adb7e3c5d345a90f34f848f6a1543309d5acc7fb78b28c4 config.json
5
  e9f179aad7619153724920c86a90ecc3c161fd15ffc3332f200e9f2073a07f72 configuration_bashkir_roberta.py
6
  48e4d4e7b8af675d222e58e14776abac24b7e83a937d7884bd9d735ea9dd2173 model.safetensors