Instructions to use mutaician/kalepres with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mutaician/kalepres with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
KalePres
KalePres is a binary Kalenjin presence detector. It determines whether Kalenjin occurs anywhere in a piece of text, including a short Kalenjin span inside English, Kiswahili, Sheng or other surrounding text.
Conventional language identification normally assigns one dominant language to an entire sentence or document. KalePres is designed for a different question:
Is there identifiable Kalenjin anywhere in this text?
The model returns an item-level presence score and token-level evidence that can be converted into highlighted text spans. Its output is one broad Kalenjin target; it does not produce dialect labels.
Highlights
- Detects full Kalenjin text and Kalenjin embedded in multilingual text.
- Trained directly for sparse presence, including one-to-three-word spans.
- Returns evidence scores for the parts of the input that triggered detection.
- Uses one inclusive Kalenjin presence target instead of requiring a specific dialect prediction.
- Uses a parameter-efficient LoRA adaptation of the AfroScope encoder.
- Outperforms whole-line and overlapping-window AfroScope and GlotLID baselines on the controlled presence benchmark.
Model details
| Property | Value |
|---|---|
| Task | Binary Kalenjin presence detection with token evidence |
| Base encoder | UBC-NLP/afroscope-model |
| Adaptation | LoRA on attention query and value projections |
| LoRA configuration | Rank 8, alpha 16, dropout 0.05 |
| Evidence head | Binary linear token head |
| Item score | Maximum eligible token score |
| Input length | 256 subword tokens |
| Recommended threshold | 0.4225 |
| Adaptation parameters | 295,681 trainable parameters |
The recommended threshold was selected on validation data for high recall at approximately 1% observed false-positive rate. Applications can choose a different threshold when their preferred balance between recall and false positives differs.
Intended use
KalePres is intended for:
- Finding text that contains Kalenjin within multilingual collections.
- Recovering code-switched text that dominant-language identification may assign to English, Kiswahili or another language.
- Filtering and ranking candidate text for corpus preparation.
- Highlighting likely Kalenjin evidence for human review.
- Research on Kalenjin presence detection and sparse-language retrieval.
KalePres is a presence detector. A positive output means the model found text that resembles learned Kalenjin evidence; it is not a dialect-classification result.
Usage
Install the required libraries:
pip install torch transformers peft huggingface_hub
Load KalePres from mutaician/kalepres and classify a piece of text:
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from torch import nn
from transformers import AutoModel, AutoTokenizer
MODEL_ID = "mutaician/kalepres"
THRESHOLD = 0.4225
model_dir = Path(snapshot_download(MODEL_ID))
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
saved_head = torch.load(
model_dir / "evidence_head.pt",
map_location="cpu",
weights_only=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_dir / "tokenizer", use_fast=True)
base = AutoModel.from_pretrained(
saved_head["base_model"],
add_pooling_layer=False,
)
encoder = PeftModel.from_pretrained(base, model_dir / "adapter").to(device).eval()
evidence_head = nn.Linear(int(saved_head["hidden_size"]), 1)
evidence_head.load_state_dict(saved_head["state_dict"])
evidence_head = evidence_head.to(device).eval()
text = (
"I remembered a kipsigis childhood song; Chepii wee, oinon teta kotapala mi sang, ak inkokiet"
)
encoded = tokenizer(
text,
truncation=True,
max_length=256,
return_offsets_mapping=True,
return_tensors="pt",
)
offsets = encoded.pop("offset_mapping")[0]
eligible = offsets[:, 0] != offsets[:, 1]
inputs = {key: value.to(device) for key, value in encoded.items()}
with torch.inference_mode():
hidden = encoder(**inputs).last_hidden_state
token_scores = torch.sigmoid(evidence_head(hidden).squeeze(-1))[0].cpu()
score = float(token_scores[eligible].max()) if eligible.any() else 0.0
prediction = "KALENJIN_PRESENT" if score >= THRESHOLD else "NOT_DETECTED"
print({"prediction": prediction, "score": score})
The output score is a detection score rather than a calibrated probability. Use the included threshold as the default operating point.
Example result
KALENJIN_PRESENT score=1.0000 decision_threshold=0.4225
Evidence:
[40:93] 'Chepii wee, oinon teta, kotapala mi sang, ak inkokiet' token_score=1.0000
Training data
Training material was adapted primarily from Anv language datasets, together with supplementary Kalenjin-labelled text. The source material was normalized, deduplicated and converted into the binary presence format required by the model. Training examples included complete Kalenjin text, mixed-language clauses, sparse one-to-three-word presence, ordinary negatives and challenging negative examples.
All included Kalenjin varieties were mapped to the single kln presence
target. Variety information was used as source metadata rather than as a model
output.
Evaluation
Full controlled test
The full controlled evaluation contains 47,203 examples. At the frozen
threshold of 0.4225, KalePres produced:
| Metric | Result |
|---|---|
| Precision | 99.23% |
| Recall | 98.71% |
| F1 | 98.96% |
| False-positive rate | 1.04% |
| Unmixed Kalenjin recall | 99.81% |
| Mixed-clause recall | 98.93% |
| Sparse 1–3-word recall | 96.86% |
Comparison with language identification
The following models were evaluated on the same fixed 10,000-example controlled presence benchmark. Windowed LID uses overlapping eight-word windows and counts an item as positive when any window receives a Kalenjin member-language label.
| Model | Method | Precision | Recall | F1 | FPR |
|---|---|---|---|---|---|
| KalePres | Token evidence | 99.31% | 98.18% | 98.74% | 1.03% |
| AfroScope | Whole line | 100.00% | 46.57% | 63.54% | 0.00% |
| AfroScope | 8-word windows | 99.65% | 61.62% | 76.15% | 0.33% |
| GlotLID v1 | Whole line | 100.00% | 42.48% | 59.63% | 0.00% |
| GlotLID v1 | 8-word windows | 99.77% | 57.38% | 72.86% | 0.20% |
| GlotLID v3 | Whole line | 100.00% | 40.78% | 57.94% | 0.00% |
| GlotLID v3 | 8-word windows | 99.58% | 55.80% | 71.52% | 0.35% |
The largest difference appears in the sparse one-to-three-word slice:
| Model | Sparse recall |
|---|---|
| KalePres | 94.67% |
| AfroScope, 8-word windows | 6.87% |
| GlotLID v1, 8-word windows | 2.73% |
| GlotLID v3, 8-word windows | 2.67% |
These results demonstrate the difference between dominant-language identification and detecting whether a target language occurs anywhere in the input.
Limitations
- The reported benchmarks use controlled mixtures; natural-world generalization has not yet been established. Further evaluations are being conducted.
- Kalenjin-looking invented or meaningless text can receive a positive score. KalePres detects learned language patterns and does not certify grammatical correctness or meaning.
- Personal names and place names can sometimes trigger positive predictions, especially when presented without enough surrounding context.
License
KalePres is released under the
Creative Commons Attribution 4.0 International
license (CC BY 4.0).
- Downloads last month
- -
Model tree for mutaician/kalepres
Base model
UBC-NLP/afroscope-model