Latensis RTE — Turkish Recognizing Textual Entailment
Understanding Beyond the Visible
Latensis RTE is a Turkish textual entailment model fine-tuned on top of Latensis RoBERTa Base using a two-stage training pipeline: NLI pretraining → RTE fine-tuning.
Labels
| ID | Label |
|---|---|
| 0 | not_entailment |
| 1 | entailment |
Benchmark Results
Evaluated on TrGLUE RTE test set (1,000 examples).
| Metric | Our Model | BERTurk NLI |
|---|---|---|
| Accuracy | 0.7960 | 0.7780 |
| F1-macro | 0.7959 | 0.7770 |
BERTurk NLI model was trained on 482k NLI examples. Our model achieves better results with only 50k NLI examples (~10x less data).
Training Details
| Property | Value |
|---|---|
| Base model | Latensis RoBERTa Base (500k steps, val loss 3.21) |
| Stage 1 | NLI pretraining — 50k examples (emrecan/all-nli-tr) |
| Stage 2 | RTE fine-tuning — 3,784 examples (TrGLUE RTE) |
| Tokenizer | Hecemen Unigram 128k |
| Learning rate | 4e-6 |
| Batch size | 16 |
| Epochs | 8 (best at epoch 7) |
Usage
import torch
import sentencepiece as spm
from transformers import RobertaForSequenceClassification, RobertaConfig
from huggingface_hub import hf_hub_download
# Load tokenizer
spm_path = hf_hub_download(
repo_id="mursideaki/hecemen-tokenizer-unigram-128k",
filename="tr_unigram_tokenizer.model"
)
sp = spm.SentencePieceProcessor()
sp.load(spm_path)
PAD_ID = sp.piece_to_id("<pad>")
BOS_ID = sp.piece_to_id("<s>")
EOS_ID = sp.piece_to_id("</s>")
MAX_LENGTH = 256
# Load model
config = RobertaConfig.from_pretrained("mursideaki/latensis-rte-tr")
model = RobertaForSequenceClassification.from_pretrained(
"mursideaki/latensis-rte-tr", config=config
)
model.eval()
def predict_rte(sentence1, sentence2):
ids = sp.encode_as_ids(sentence1)
ids += [EOS_ID]
ids += sp.encode_as_ids(sentence2)
ids = [BOS_ID] + ids[:MAX_LENGTH-2] + [EOS_ID]
mask = [1] * len(ids)
if len(ids) < MAX_LENGTH:
pad_len = MAX_LENGTH - len(ids)
ids = ids + [PAD_ID] * pad_len
mask = mask + [0] * pad_len
with torch.no_grad():
outputs = model(
input_ids=torch.tensor([ids], dtype=torch.long),
attention_mask=torch.tensor([mask], dtype=torch.long)
)
label = outputs.logits.argmax(dim=-1).item()
return "entailment" if label == 1 else "not_entailment"
print(predict_rte(
"Türkiye'nin başkenti Ankara'dır.",
"Ankara, Türkiye'de bir şehirdir."
))
# → entailment
Companion Models
- latensis-roberta-base-tr — Base model
- latensis-sentiment-tr — Sentiment
- latensis-ner-tr — NER
- latensis-sts-tr — STS
- hecemen-tokenizer-unigram-128k — Tokenizer
Citation
@misc{latensis2026,
author = {Mürşide Aki},
title = {Latensis: Turkish NLP Model Suite},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/mursideaki/latensis-rte-tr}
}
License
MIT
- Downloads last month
- 21