laya-fa-support
Fine-tuned by Ali Jahani · Telegram: t.me/tarfandoonchannel
A Persian customer-support triage model, fine-tuned from
Laya's laya-multilingual checkpoint. It reads a
Persian support message and decides which department should handle it (billing,
technical, cancel or other) in a single forward pass. It returns a probability for every
option and generates no text, so there is nothing to parse and nothing to hallucinate.
Results at a glance
These scores come from the held-out persian_fa benchmark: 64 Persian
support messages the model never saw during training, asked with prompts it never saw
during training.
| Prompt |
Base laya-multilingual |
This model |
Change |
| Persian instructions |
36/64 (56.2%) |
51/64 (79.7%) |
+15 cases |
| English instructions |
41/64 (64.1%) |
44/64 (68.8%) |
+3 cases |
Quick start
pip install laya
import laya
agent = laya.load("tarfandoon/laya-fa-support")
question = {
"department": {
"type": "choice",
"instructions": "این پیام پشتیبانی باید به کدام بخش ارجاع داده شود؟",
"criteria": {
"billing": "پرداخت، کسر وجه، استرداد پول، فاکتور، قیمت و هزینهی اشتراک",
"technical": "خطا، باگ، کرش، مشکل ورود یا رمز، کندی و قطعی برنامه",
"cancel": "کاربر میخواهد اشتراک را لغو کند، تمدید را متوقف کند یا حساب را ببندد یا حذف کند",
"other": "هر چیز دیگر: تشکر، پیشنهاد، همکاری، استخدام، سؤال عمومی",
},
}
}
result = agent.predict({"message": "سلام، دو بار پول از حسابم کم شده ولی اشتراکم فعال نشده"}, question)
answer = result["answers"]["department"]
print(answer["choice"])
print(answer["probabilities"])
print(answer["answer_confidence"])
Many messages at once
messages = ["اپ باز نمیشه", "mikham eshterakam ro laghv konam", "مرسی از پشتیبانیتون"]
results = agent.predict_batch([{"message": m} for m in messages], question)
for m, r in zip(messages, results):
print(r["answers"]["department"]["choice"], "<-", m)
Send uncertain cases to a human
THRESHOLD = 0.6
if answer["answer_confidence"] >= THRESHOLD:
route_to(answer["choice"])
else:
send_to_human(answer)
Tips
- Keep the four option keys (
billing, technical, cancel, other) and give every
option a full description, like the example above. The wording can change, because the
model was trained on several paraphrases in Persian and English. Persian instructions
work best for this model (see the results).
- ⚠️ Do not use one-word descriptions such as
"billing": "مالی". In a spot check on 8 new
messages, the full prompt above got 7/8 right, while one-word descriptions got 5/8. The
one-word version was also confidently wrong: it labelled «اپ باز نمیشه» as cancel with
0.92 confidence.
- The input can be formal or colloquial Persian, Finglish, or Persian mixed with English words.
Arabic ي/ك and missing half-spaces are handled.
- On a CPU a decision takes roughly 150-200 ms. On a GPU it is much faster.
Full benchmark results: before vs after
Base model: convaiinnovations/laya / multilingual @ 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Scores are from the first
of three repeats, and every repeat gave identical predictions. Neither model saw these 64
messages or these two prompts during training.
Persian instructions (choice_fa)
| Metric |
Base |
This model |
Change |
| Accuracy |
36/64 |
51/64 |
+15 |
| Macro-F1 |
0.513 |
0.795 |
|
False cancel (would wrongly close an account) |
4 |
2 |
|
Missed cancel |
9 |
3 |
|
| Mean answer confidence |
0.84 |
0.69 |
|
| ECE (calibration error, lower is better) |
0.273 |
0.103 |
|
| Family |
Base |
This model |
Change |
| formal |
6/8 |
8/8 |
+2 |
| colloquial |
5/8 |
7/8 |
+2 |
| finglish |
4/8 |
4/8 |
0 |
| orthography |
4/8 |
6/8 |
+2 |
| code_mixed |
5/8 |
8/8 |
+3 |
| negation |
3/8 |
6/8 |
+3 |
| sarcasm_taarof |
5/8 |
6/8 |
+1 |
| digits |
4/8 |
6/8 |
+2 |
| Label (recall) |
Base |
This model |
billing |
13/16 |
14/16 |
technical |
14/16 |
14/16 |
cancel |
7/16 |
13/16 |
other |
2/16 |
10/16 |
Confusion matrices (rows = gold, columns = predicted)
Base
|
billing |
technical |
cancel |
other |
| billing |
13 |
2 |
1 |
0 |
| technical |
1 |
14 |
1 |
0 |
| cancel |
1 |
8 |
7 |
0 |
| other |
6 |
6 |
2 |
2 |
This model
|
billing |
technical |
cancel |
other |
| billing |
14 |
0 |
1 |
1 |
| technical |
1 |
14 |
1 |
0 |
| cancel |
1 |
2 |
13 |
0 |
| other |
4 |
2 |
0 |
10 |
English instructions (choice_en)
| Metric |
Base |
This model |
Change |
| Accuracy |
41/64 |
44/64 |
+3 |
| Macro-F1 |
0.633 |
0.685 |
|
False cancel (would wrongly close an account) |
2 |
3 |
|
Missed cancel |
6 |
6 |
|
| Mean answer confidence |
0.82 |
0.71 |
|
| ECE (calibration error, lower is better) |
0.250 |
0.108 |
|
| Family |
Base |
This model |
Change |
| formal |
5/8 |
8/8 |
+3 |
| colloquial |
7/8 |
6/8 |
-1 |
| finglish |
4/8 |
5/8 |
+1 |
| orthography |
5/8 |
5/8 |
0 |
| code_mixed |
6/8 |
6/8 |
0 |
| negation |
3/8 |
4/8 |
+1 |
| sarcasm_taarof |
6/8 |
5/8 |
-1 |
| digits |
5/8 |
5/8 |
0 |
| Label (recall) |
Base |
This model |
billing |
10/16 |
10/16 |
technical |
15/16 |
14/16 |
cancel |
10/16 |
10/16 |
other |
6/16 |
10/16 |
Confusion matrices (rows = gold, columns = predicted)
Base
|
billing |
technical |
cancel |
other |
| billing |
10 |
3 |
2 |
1 |
| technical |
1 |
15 |
0 |
0 |
| cancel |
0 |
6 |
10 |
0 |
| other |
4 |
6 |
0 |
6 |
This model
|
billing |
technical |
cancel |
other |
| billing |
10 |
2 |
2 |
2 |
| technical |
1 |
14 |
1 |
0 |
| cancel |
0 |
4 |
10 |
2 |
| other |
3 |
3 |
0 |
10 |
Training
|
|
| Base checkpoint |
laya-multilingual (mmBERT-base, 322M parameters) @ 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 |
| Method |
top mode: decision head + last 4 encoder layers. Token embeddings frozen. |
| Objective |
cross-entropy on the gold option, AdamW, fp16 autocast |
| Data |
200 training and 40 dev Persian support messages (50 and 10 per label), all synthetic |
| Prompts |
4 paraphrased questions (2 Persian, 2 English), none of them the benchmark prompts |
| Robustness |
options shuffled with probability 0.5; random ي/ك, half-space and digit variants |
| Model selection |
best epoch by dev NLL (epoch 6) |
| Calibration |
one choice temperature fitted on dev (1.27) |
| Hardware |
a single NVIDIA GeForce GTX 1650 (4 GB), about 10 minutes |
| Software |
laya 0.3.21, torch 2.14.0+cu126, seed 42 |
The dev set has 40 messages x 4 training prompts. Dev accuracy went from 0.562
before training to 0.700 at the selected epoch.
| Epoch |
Train loss |
Dev accuracy |
Dev NLL |
| 0 (before) |
- |
0.562 |
1.555 |
| 1 |
1.077 |
0.581 |
0.999 |
| 2 |
0.870 |
0.619 |
0.917 |
| 3 |
0.800 |
0.656 |
0.896 |
| 4 |
0.800 |
0.663 |
0.852 |
| 5 |
0.728 |
0.700 |
0.839 |
| 6 ✅ selected |
0.733 |
0.700 |
0.810 |
| 7 |
0.700 |
0.706 |
0.818 |
| 8 |
0.639 |
0.706 |
0.820 |
A near-duplicate check (character 3-gram Jaccard < 0.5) guarantees that no training or dev
message copies a benchmark message.
Limitations
- The data is small and synthetic. The training set and the benchmark were drafted the
same way, with AI assistance. Real customer traffic will be harder, so expect a smaller
gain there than the benchmark shows. Evaluate on your own messages before relying on the
model.
- With Persian instructions, these families did not improve: finglish (4/8).
- With English instructions, these families did not improve: colloquial (6/8), orthography (5/8), code_mixed (6/8), sarcasm_taarof (5/8), digits (5/8).
- With English instructions the overall change is small (+3 of 64), which is within noise for 64 cases.
- It is sensitive to how the options are described. Use descriptive option texts (see
Tips). Terse one-word options made it much worse, with high confidence.
- It is specialised to one 4-way decision. Other Laya tasks and languages were not
re-evaluated and may have degraded. Use the original Laya checkpoints for those.
- Confidence was calibrated on 160 dev rows only. Refit the threshold or temperature on your
own data before gating decisions on it.
Links
License
Apache-2.0, the same as the base Laya weights.