KalePres

KalePres is a binary Kalenjin presence detector. It determines whether Kalenjin occurs anywhere in a piece of text, including a short Kalenjin span inside English, Kiswahili, Sheng or other surrounding text.

Conventional language identification normally assigns one dominant language to an entire sentence or document. KalePres is designed for a different question:

Is there identifiable Kalenjin anywhere in this text?

The model returns an item-level presence score and token-level evidence that can be converted into highlighted text spans. Its output is one broad Kalenjin target; it does not produce dialect labels.

Highlights

  • Detects full Kalenjin text and Kalenjin embedded in multilingual text.
  • Trained directly for sparse presence, including one-to-three-word spans.
  • Returns evidence scores for the parts of the input that triggered detection.
  • Uses one inclusive Kalenjin presence target instead of requiring a specific dialect prediction.
  • Uses a parameter-efficient LoRA adaptation of the AfroScope encoder.
  • Outperforms whole-line and overlapping-window AfroScope and GlotLID baselines on the controlled presence benchmark.

Model details

Property Value
Task Binary Kalenjin presence detection with token evidence
Base encoder UBC-NLP/afroscope-model
Adaptation LoRA on attention query and value projections
LoRA configuration Rank 8, alpha 16, dropout 0.05
Evidence head Binary linear token head
Item score Maximum eligible token score
Input length 256 subword tokens
Recommended threshold 0.4225
Adaptation parameters 295,681 trainable parameters

The recommended threshold was selected on validation data for high recall at approximately 1% observed false-positive rate. Applications can choose a different threshold when their preferred balance between recall and false positives differs.

Intended use

KalePres is intended for:

  • Finding text that contains Kalenjin within multilingual collections.
  • Recovering code-switched text that dominant-language identification may assign to English, Kiswahili or another language.
  • Filtering and ranking candidate text for corpus preparation.
  • Highlighting likely Kalenjin evidence for human review.
  • Research on Kalenjin presence detection and sparse-language retrieval.

KalePres is a presence detector. A positive output means the model found text that resembles learned Kalenjin evidence; it is not a dialect-classification result.

Usage

Install the required libraries:

pip install torch transformers peft huggingface_hub

Load KalePres from mutaician/kalepres and classify a piece of text:

from pathlib import Path

import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from torch import nn
from transformers import AutoModel, AutoTokenizer


MODEL_ID = "mutaician/kalepres"
THRESHOLD = 0.4225

model_dir = Path(snapshot_download(MODEL_ID))
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

saved_head = torch.load(
    model_dir / "evidence_head.pt",
    map_location="cpu",
    weights_only=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_dir / "tokenizer", use_fast=True)
base = AutoModel.from_pretrained(
    saved_head["base_model"],
    add_pooling_layer=False,
)
encoder = PeftModel.from_pretrained(base, model_dir / "adapter").to(device).eval()
evidence_head = nn.Linear(int(saved_head["hidden_size"]), 1)
evidence_head.load_state_dict(saved_head["state_dict"])
evidence_head = evidence_head.to(device).eval()

text = (
    "I remembered a kipsigis childhood song; Chepii wee, oinon teta kotapala mi sang, ak inkokiet"
)
encoded = tokenizer(
    text,
    truncation=True,
    max_length=256,
    return_offsets_mapping=True,
    return_tensors="pt",
)
offsets = encoded.pop("offset_mapping")[0]
eligible = offsets[:, 0] != offsets[:, 1]
inputs = {key: value.to(device) for key, value in encoded.items()}

with torch.inference_mode():
    hidden = encoder(**inputs).last_hidden_state
    token_scores = torch.sigmoid(evidence_head(hidden).squeeze(-1))[0].cpu()

score = float(token_scores[eligible].max()) if eligible.any() else 0.0
prediction = "KALENJIN_PRESENT" if score >= THRESHOLD else "NOT_DETECTED"
print({"prediction": prediction, "score": score})

The output score is a detection score rather than a calibrated probability. Use the included threshold as the default operating point.

Example result

KALENJIN_PRESENT  score=1.0000  decision_threshold=0.4225
Evidence:
  [40:93] 'Chepii wee, oinon teta, kotapala mi sang, ak inkokiet'  token_score=1.0000

Training data

Training material was adapted primarily from Anv language datasets, together with supplementary Kalenjin-labelled text. The source material was normalized, deduplicated and converted into the binary presence format required by the model. Training examples included complete Kalenjin text, mixed-language clauses, sparse one-to-three-word presence, ordinary negatives and challenging negative examples.

All included Kalenjin varieties were mapped to the single kln presence target. Variety information was used as source metadata rather than as a model output.

Evaluation

Full controlled test

The full controlled evaluation contains 47,203 examples. At the frozen threshold of 0.4225, KalePres produced:

Metric Result
Precision 99.23%
Recall 98.71%
F1 98.96%
False-positive rate 1.04%
Unmixed Kalenjin recall 99.81%
Mixed-clause recall 98.93%
Sparse 1–3-word recall 96.86%

Comparison with language identification

The following models were evaluated on the same fixed 10,000-example controlled presence benchmark. Windowed LID uses overlapping eight-word windows and counts an item as positive when any window receives a Kalenjin member-language label.

Model Method Precision Recall F1 FPR
KalePres Token evidence 99.31% 98.18% 98.74% 1.03%
AfroScope Whole line 100.00% 46.57% 63.54% 0.00%
AfroScope 8-word windows 99.65% 61.62% 76.15% 0.33%
GlotLID v1 Whole line 100.00% 42.48% 59.63% 0.00%
GlotLID v1 8-word windows 99.77% 57.38% 72.86% 0.20%
GlotLID v3 Whole line 100.00% 40.78% 57.94% 0.00%
GlotLID v3 8-word windows 99.58% 55.80% 71.52% 0.35%

The largest difference appears in the sparse one-to-three-word slice:

Model Sparse recall
KalePres 94.67%
AfroScope, 8-word windows 6.87%
GlotLID v1, 8-word windows 2.73%
GlotLID v3, 8-word windows 2.67%

These results demonstrate the difference between dominant-language identification and detecting whether a target language occurs anywhere in the input.

Limitations

  • The reported benchmarks use controlled mixtures; natural-world generalization has not yet been established. Further evaluations are being conducted.
  • Kalenjin-looking invented or meaningless text can receive a positive score. KalePres detects learned language patterns and does not certify grammatical correctness or meaning.
  • Personal names and place names can sometimes trigger positive predictions, especially when presented without enough surrounding context.

License

KalePres is released under the Creative Commons Attribution 4.0 International license (CC BY 4.0).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mutaician/kalepres

Adapter
(1)
this model