Latensis RTE — Turkish Recognizing Textual Entailment

Understanding Beyond the Visible

Latensis RTE is a Turkish textual entailment model fine-tuned on top of Latensis RoBERTa Base using a two-stage training pipeline: NLI pretraining → RTE fine-tuning.

Labels

ID Label
0 not_entailment
1 entailment

Benchmark Results

Evaluated on TrGLUE RTE test set (1,000 examples).

Metric Our Model BERTurk NLI
Accuracy 0.7960 0.7780
F1-macro 0.7959 0.7770

BERTurk NLI model was trained on 482k NLI examples. Our model achieves better results with only 50k NLI examples (~10x less data).

Training Details

Property Value
Base model Latensis RoBERTa Base (500k steps, val loss 3.21)
Stage 1 NLI pretraining — 50k examples (emrecan/all-nli-tr)
Stage 2 RTE fine-tuning — 3,784 examples (TrGLUE RTE)
Tokenizer Hecemen Unigram 128k
Learning rate 4e-6
Batch size 16
Epochs 8 (best at epoch 7)

Usage

import torch
import sentencepiece as spm
from transformers import RobertaForSequenceClassification, RobertaConfig
from huggingface_hub import hf_hub_download

# Load tokenizer
spm_path = hf_hub_download(
    repo_id="mursideaki/hecemen-tokenizer-unigram-128k",
    filename="tr_unigram_tokenizer.model"
)
sp = spm.SentencePieceProcessor()
sp.load(spm_path)

PAD_ID = sp.piece_to_id("<pad>")
BOS_ID = sp.piece_to_id("<s>")
EOS_ID = sp.piece_to_id("</s>")
MAX_LENGTH = 256

# Load model
config = RobertaConfig.from_pretrained("mursideaki/latensis-rte-tr")
model  = RobertaForSequenceClassification.from_pretrained(
    "mursideaki/latensis-rte-tr", config=config
)
model.eval()

def predict_rte(sentence1, sentence2):
    ids  = sp.encode_as_ids(sentence1)
    ids += [EOS_ID]
    ids += sp.encode_as_ids(sentence2)
    ids  = [BOS_ID] + ids[:MAX_LENGTH-2] + [EOS_ID]
    mask = [1] * len(ids)
    if len(ids) < MAX_LENGTH:
        pad_len = MAX_LENGTH - len(ids)
        ids  = ids  + [PAD_ID] * pad_len
        mask = mask + [0]      * pad_len

    with torch.no_grad():
        outputs = model(
            input_ids=torch.tensor([ids], dtype=torch.long),
            attention_mask=torch.tensor([mask], dtype=torch.long)
        )
    label = outputs.logits.argmax(dim=-1).item()
    return "entailment" if label == 1 else "not_entailment"

print(predict_rte(
    "Türkiye'nin başkenti Ankara'dır.",
    "Ankara, Türkiye'de bir şehirdir."
))
# → entailment

Companion Models

Citation

@misc{latensis2026,
  author    = {Mürşide Aki},
  title     = {Latensis: Turkish NLP Model Suite},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/mursideaki/latensis-rte-tr}
}

License

MIT

Downloads last month
21
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support