Instructions to use kubataba/ru-kbd-bidirectional with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kubataba/ru-kbd-bidirectional with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="kubataba/ru-kbd-bidirectional")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("kubataba/ru-kbd-bidirectional") model = AutoModelForSeq2SeqLM.from_pretrained("kubataba/ru-kbd-bidirectional", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Russian ↔ Kabardian bidirectional translator (MarianMT, 61M)
One small model translates in both directions between Russian and Kabardian (East Circassian, kbd). The
direction is chosen by a target tag at the start of the source text: >>kbd<< (into Kabardian) or >>ru<< (into
Russian).
It was trained from our earlier one-direction model kubataba/ru-kbd-opus
with a new tokenizer and five rounds of training with iterative back-translation. On FLORES-200 it reaches
chrF 55.9 for ru→kbd and 49.6 for kbd→ru, against 38.7 / 35.0 for our earlier one-direction models.
| architecture | MarianMT, 6 + 6 layers, d_model 512, 8 heads, 61M parameters |
| tokenizer | SentencePiece unigram, Russian + Kabardian, 32 001 pieces + 44 appended characters = 32 045 |
| palochka | always ӏ U+04CF (not Latin I, not Ӏ U+04C0) — normalize your input |
| decoding | beam 4 (greedy: −1.1 / −1.0 chrF, about 2× faster) |
| licence | CC-BY-NC-4.0 |
Usage
from transformers import MarianMTModel, MarianTokenizer
name = "kubataba/ru-kbd-bidirectional"
tok = MarianTokenizer.from_pretrained(name)
model = MarianMTModel.from_pretrained(name).eval()
def translate(text, target): # target: "kbd" or "ru"
batch = tok([f">>{target}<< {text}"], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_new_tokens=160)
return tok.decode(out[0], skip_special_tokens=True)
print(translate("Сейчас он работает в Москве.", "kbd")) # Иджыпсту ар Москва щолажьэ
print(translate("Иджы псори нэгъуэщӏущ.", "ru")) # Теперь все иначе.
Write the Kabardian palochka as ӏ (U+04CF). Inputs with Latin I or 1 in its place should be normalized first.
Recommended use: translate sentence by sentence (the training pairs average 7 words; whole paragraphs lose content), and stop the generation if a short piece repeats four times in a row (degenerate loops are rare but possible, mostly in the int8 version).
Results
FLORES-200 (devtest, first 200 sentences, Russian side from FLORES). The 200 Kabardian references were made for this evaluation with Yandex Translate: sentences 1–50 were checked and corrected by a native speaker, sentences 51–200 are the machine translation used as is. A machine-made reference favours outputs that resemble that system, so treat the 51–200 part as indicative; the 1–50 column is the human-checked one. The references are not distributed. FLORES sentences were removed from all training data.
| model | ru→kbd (200) | kbd→ru (200) | ru→kbd (1–50, human-checked) | kbd→ru (1–50) |
|---|---|---|---|---|
kubataba/ru-kbd-opus / kubataba/kbd-ru-opus (beam 4) |
38.7 | 35.0 | — | — |
| this model, beam 4 | 55.9 | 49.6 | 56.4 | 50.1 |
| this model, int8 ONNX, beam 4 | 55.7 | 49.6 | — | — |
| this model, int8 ONNX, greedy | 54.6 | 48.6 | — | — |
English → Kabardian through Russian (FLORES, English source; en→ru by MADLAD-400 3B, then this model): chrF 46.6 with beam 4, 45.7 greedy. The reference is the Kabardian translation of the human Russian sentence, so a cascade that words the Russian differently is underestimated.
anzorq test set (1 081 pairs, mostly dictionary entries — very short): ru→kbd 45.8, kbd→ru 40.5. Short isolated words are the hardest case for this model; it is trained mostly on sentences.
Latin words in the source are carried through to the output in 98% of cases (46 Latin words in FLORES).
Training
Tokenizer. The opus-mt tokenizers of the base models split Kabardian almost into letters (5.5–7.6 tokens per
word). We trained a joint Russian–Kabardian SentencePiece unigram tokenizer (1.7 tokens per Kabardian word, 1.5 per
Russian). The transformer body of kubataba/ru-kbd-opus was kept; each new piece's embedding started as the mean of
the old embeddings of the pieces the old tokenizer split it into. At stage 3 we appended 44 characters the corpus
had never shown (the hyphen, Latin letters, … / № $ € and others) without renumbering the vocabulary.
Parallel data. About 300 000 Russian–Kabardian pairs: our cleaned compilation from
adiga-ai/circassian-parallel-corpus and
adiga-ai/circassian-russian_texts_synthetic,
filtered: digit mismatches repaired or dropped, Adyghe (West Circassian) lines removed by a character n-gram
classifier, Russian copies on the Kabardian side removed, duplicates removed, and every pair that overlaps the test
sets or FLORES removed. Each pair is used in both directions.
Synthetic data — the target side is always real text. Machine-made text is only ever on the input side, so the model learns to produce real Kabardian and real Russian.
kubataba/kbd-ru-synthetic-backtranslation— 591 726 real Kabardian sentences with machine Russian; teaches ru→kbd.kubataba/ru-kbd-synthetic-forward-translation— 587 913 real Russian sentences with machine Kabardian; teaches kbd→ru.
This model was trained on version 1 of both corpora (config "v1", tag v1). Version 2 (the default since
05.10.2026) removes 7 426 and 2 203 rows, translates 16 735 rows of the forward-translation corpus again and corrects
digits and 26 calques; see the dataset cards.
Every synthetic pair passes a filter: identical sides, copies (character 5-gram Jaccard ≥ 0.5), palochka in the Russian output, truncated output, and degenerate loops (a 1–6 word span three times in a row that is not in the source) are dropped.
Five stages (RTX 4090, about 6 hours in total; FLORES chrF with beam 4):
| stage | what changed | ru→kbd | kbd→ru |
|---|---|---|---|
| 1 | new tokenizer; parallel corpus both ways + back-translation of 613k Kabardian sentences by kubataba/kbd-ru-opus; 30k steps |
47.6 | 35.7 |
| 2 | + forward translation of 575k Russian sentences by the stage-1 model; 20k steps | 49.0 | 45.1 |
| 3 | + 44 characters in the vocabulary, + 20k pairs carrying Latin words; 10k steps | 50.5 | 47.2 |
| 4 | iterative back-translation: the Kabardian sentences translated again by the stage-3 model; 15k steps | 55.1 | 47.9 |
| 5 | the Russian sentences translated again by the stage-4 model; 15k steps | 55.9 | 49.6 |
Each re-translation by a stronger model gave better synthetic data (stage 4: hyphens kept in 51 500 sentences instead of 115, degenerate loops 676 instead of 4 985). Training stopped when a further round was predicted to add less than one point. Checkpoint averaging and beams of 6 or 8 did not improve on beam 4.
Hyperparameters: AdamW, weight decay 0.01, label smoothing 0.1, inverse square-root schedule with 5% warm-up, batch 128, max length 128 tokens; lr 3e-4 (stage 1, after 1 000 steps training only the new embeddings), 1.5e-4 (stage 2), 1e-4 (stages 3–5); bf16 on the GPU; the best checkpoint by FLORES dev chrF (greedy) every 1 000 steps.
Limitations
- Quoted text. In the Kabardian press, names in quotes stay in Russian, and the model learned to copy quoted text. For direct speech in quotes, translate the inside of the quotes separately and put the quotes back (in our app: quoted spans of 4+ words are translated apart; shorter ones lose the quotes and are translated in context). This lowers the number of FLORES sentences with a copied Russian word from 36 to 21.
- Rare words (technical, archaic) may be left in Russian; a dictionary substitution is not a fix — Kabardian needs the inflected form.
- Long sentences (> 40 words) are rare in training; split at clause boundaries.
- The English→Kabardian quality depends on the English→Russian step.
Licence and attribution
The model is released under CC-BY-NC-4.0 (non-commercial). The copyright holder uses it in the SayFable app.
- Base model:
kubataba/ru-kbd-opus(CC-BY-NC-4.0), itself trained from Helsinki-NLP OPUS-MT (Apache-2.0). - Parallel data: Anzor Qunash (adiga.ai), Circassian–Russian Parallel Corpus (CC-BY-4.0) and Circassian–Russian texts (MIT).
- Monolingual Kabardian:
anzorq/kbd_monolingual(MIT), texts of adyghepsale.ru, apkbr.ru, adygabza.ru, cherkes-haky.ru. - Monolingual Russian: Russian Wikipedia and Wikisource (CC-BY-SA); 19th-century Russian fiction from
nevmenandr/accentual-syllabic-verse-in-russian-prose(MIT; B. V. Orekhov, 462 texts). - Evaluation: FLORES-200 (CC-BY-SA-4.0); English→Russian in the cascade: MADLAD-400 3B (Apache-2.0).
- Downloads last month
- 18