Russian ↔ Kabardian bidirectional translator (MarianMT, 61M)

One small model translates in both directions between Russian and Kabardian (East Circassian, kbd). The direction is chosen by a target tag at the start of the source text: >>kbd<< (into Kabardian) or >>ru<< (into Russian).

It was trained from our earlier one-direction model kubataba/ru-kbd-opus with a new tokenizer and five rounds of training with iterative back-translation. On FLORES-200 it reaches chrF 55.9 for ru→kbd and 49.6 for kbd→ru, against 38.7 / 35.0 for our earlier one-direction models.

architecture MarianMT, 6 + 6 layers, d_model 512, 8 heads, 61M parameters
tokenizer SentencePiece unigram, Russian + Kabardian, 32 001 pieces + 44 appended characters = 32 045
palochka always ӏ U+04CF (not Latin I, not Ӏ U+04C0) — normalize your input
decoding beam 4 (greedy: −1.1 / −1.0 chrF, about 2× faster)
licence CC-BY-NC-4.0

Usage

from transformers import MarianMTModel, MarianTokenizer

name = "kubataba/ru-kbd-bidirectional"
tok = MarianTokenizer.from_pretrained(name)
model = MarianMTModel.from_pretrained(name).eval()

def translate(text, target):              # target: "kbd" or "ru"
    batch = tok([f">>{target}<< {text}"], return_tensors="pt")
    out = model.generate(**batch, num_beams=4, max_new_tokens=160)
    return tok.decode(out[0], skip_special_tokens=True)

print(translate("Сейчас он работает в Москве.", "kbd"))   # Иджыпсту ар Москва щолажьэ
print(translate("Иджы псори нэгъуэщӏущ.", "ru"))       # Теперь все иначе.

Write the Kabardian palochka as ӏ (U+04CF). Inputs with Latin I or 1 in its place should be normalized first.

Recommended use: translate sentence by sentence (the training pairs average 7 words; whole paragraphs lose content), and stop the generation if a short piece repeats four times in a row (degenerate loops are rare but possible, mostly in the int8 version).

Results

FLORES-200 (devtest, first 200 sentences, Russian side from FLORES). The 200 Kabardian references were made for this evaluation with Yandex Translate: sentences 1–50 were checked and corrected by a native speaker, sentences 51–200 are the machine translation used as is. A machine-made reference favours outputs that resemble that system, so treat the 51–200 part as indicative; the 1–50 column is the human-checked one. The references are not distributed. FLORES sentences were removed from all training data.

model ru→kbd (200) kbd→ru (200) ru→kbd (1–50, human-checked) kbd→ru (1–50)
kubataba/ru-kbd-opus / kubataba/kbd-ru-opus (beam 4) 38.7 35.0 — —
this model, beam 4 55.9 49.6 56.4 50.1
this model, int8 ONNX, beam 4 55.7 49.6 — —
this model, int8 ONNX, greedy 54.6 48.6 — —

English → Kabardian through Russian (FLORES, English source; en→ru by MADLAD-400 3B, then this model): chrF 46.6 with beam 4, 45.7 greedy. The reference is the Kabardian translation of the human Russian sentence, so a cascade that words the Russian differently is underestimated.

anzorq test set (1 081 pairs, mostly dictionary entries — very short): ru→kbd 45.8, kbd→ru 40.5. Short isolated words are the hardest case for this model; it is trained mostly on sentences.

Latin words in the source are carried through to the output in 98% of cases (46 Latin words in FLORES).

Training

Tokenizer. The opus-mt tokenizers of the base models split Kabardian almost into letters (5.5–7.6 tokens per word). We trained a joint Russian–Kabardian SentencePiece unigram tokenizer (1.7 tokens per Kabardian word, 1.5 per Russian). The transformer body of kubataba/ru-kbd-opus was kept; each new piece's embedding started as the mean of the old embeddings of the pieces the old tokenizer split it into. At stage 3 we appended 44 characters the corpus had never shown (the hyphen, Latin letters, … / № $ € and others) without renumbering the vocabulary.

Parallel data. About 300 000 Russian–Kabardian pairs: our cleaned compilation from adiga-ai/circassian-parallel-corpus and adiga-ai/circassian-russian_texts_synthetic, filtered: digit mismatches repaired or dropped, Adyghe (West Circassian) lines removed by a character n-gram classifier, Russian copies on the Kabardian side removed, duplicates removed, and every pair that overlaps the test sets or FLORES removed. Each pair is used in both directions.

Synthetic data — the target side is always real text. Machine-made text is only ever on the input side, so the model learns to produce real Kabardian and real Russian.

This model was trained on version 1 of both corpora (config "v1", tag v1). Version 2 (the default since 05.10.2026) removes 7 426 and 2 203 rows, translates 16 735 rows of the forward-translation corpus again and corrects digits and 26 calques; see the dataset cards.

Every synthetic pair passes a filter: identical sides, copies (character 5-gram Jaccard ≥ 0.5), palochka in the Russian output, truncated output, and degenerate loops (a 1–6 word span three times in a row that is not in the source) are dropped.

Five stages (RTX 4090, about 6 hours in total; FLORES chrF with beam 4):

stage what changed ru→kbd kbd→ru
1 new tokenizer; parallel corpus both ways + back-translation of 613k Kabardian sentences by kubataba/kbd-ru-opus; 30k steps 47.6 35.7
2 + forward translation of 575k Russian sentences by the stage-1 model; 20k steps 49.0 45.1
3 + 44 characters in the vocabulary, + 20k pairs carrying Latin words; 10k steps 50.5 47.2
4 iterative back-translation: the Kabardian sentences translated again by the stage-3 model; 15k steps 55.1 47.9
5 the Russian sentences translated again by the stage-4 model; 15k steps 55.9 49.6

Each re-translation by a stronger model gave better synthetic data (stage 4: hyphens kept in 51 500 sentences instead of 115, degenerate loops 676 instead of 4 985). Training stopped when a further round was predicted to add less than one point. Checkpoint averaging and beams of 6 or 8 did not improve on beam 4.

Hyperparameters: AdamW, weight decay 0.01, label smoothing 0.1, inverse square-root schedule with 5% warm-up, batch 128, max length 128 tokens; lr 3e-4 (stage 1, after 1 000 steps training only the new embeddings), 1.5e-4 (stage 2), 1e-4 (stages 3–5); bf16 on the GPU; the best checkpoint by FLORES dev chrF (greedy) every 1 000 steps.

Limitations

  • Quoted text. In the Kabardian press, names in quotes stay in Russian, and the model learned to copy quoted text. For direct speech in quotes, translate the inside of the quotes separately and put the quotes back (in our app: quoted spans of 4+ words are translated apart; shorter ones lose the quotes and are translated in context). This lowers the number of FLORES sentences with a copied Russian word from 36 to 21.
  • Rare words (technical, archaic) may be left in Russian; a dictionary substitution is not a fix — Kabardian needs the inflected form.
  • Long sentences (> 40 words) are rare in training; split at clause boundaries.
  • The English→Kabardian quality depends on the English→Russian step.

Licence and attribution

The model is released under CC-BY-NC-4.0 (non-commercial). The copyright holder uses it in the SayFable app.

Downloads last month
18
Safetensors
Model size
60.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kubataba/ru-kbd-bidirectional

Finetuned
(1)
this model

Datasets used to train kubataba/ru-kbd-bidirectional