laya-fa-support

Fine-tuned by Ali Jahani · Telegram: t.me/tarfandoonchannel

A Persian customer-support triage model, fine-tuned from Laya's laya-multilingual checkpoint. It reads a Persian support message and decides which department should handle it (billing, technical, cancel or other) in a single forward pass. It returns a probability for every option and generates no text, so there is nothing to parse and nothing to hallucinate.

Results at a glance

These scores come from the held-out persian_fa benchmark: 64 Persian support messages the model never saw during training, asked with prompts it never saw during training.

Prompt Base laya-multilingual This model Change
Persian instructions 36/64 (56.2%) 51/64 (79.7%) +15 cases
English instructions 41/64 (64.1%) 44/64 (68.8%) +3 cases

Quick start

pip install laya
import laya

agent = laya.load("tarfandoon/laya-fa-support")            # add device="cuda" for a GPU

question = {
    "department": {
        "type": "choice",
        "instructions": "این پیام پشتیبانی باید به کدام بخش ارجاع داده شود؟",
        "criteria": {
            "billing": "پرداخت، کسر وجه، استرداد پول، فاکتور، قیمت و هزینه‌ی اشتراک",
            "technical": "خطا، باگ، کرش، مشکل ورود یا رمز، کندی و قطعی برنامه",
            "cancel": "کاربر می‌خواهد اشتراک را لغو کند، تمدید را متوقف کند یا حساب را ببندد یا حذف کند",
            "other": "هر چیز دیگر: تشکر، پیشنهاد، همکاری، استخدام، سؤال عمومی",
        },
    }
}

result = agent.predict({"message": "سلام، دو بار پول از حسابم کم شده ولی اشتراکم فعال نشده"}, question)
answer = result["answers"]["department"]
print(answer["choice"])              # one of: billing, technical, cancel, other
print(answer["probabilities"])       # probability for every department
print(answer["answer_confidence"])   # calibrated probability of the chosen answer

Many messages at once

messages = ["اپ باز نمیشه", "mikham eshterakam ro laghv konam", "مرسی از پشتیبانی‌تون"]
results = agent.predict_batch([{"message": m} for m in messages], question)
for m, r in zip(messages, results):
    print(r["answers"]["department"]["choice"], "<-", m)

Send uncertain cases to a human

THRESHOLD = 0.6   # choose this on your own labelled data
if answer["answer_confidence"] >= THRESHOLD:
    route_to(answer["choice"])
else:
    send_to_human(answer)

Tips

  • Keep the four option keys (billing, technical, cancel, other) and give every option a full description, like the example above. The wording can change, because the model was trained on several paraphrases in Persian and English. Persian instructions work best for this model (see the results).
  • ⚠️ Do not use one-word descriptions such as "billing": "مالی". In a spot check on 8 new messages, the full prompt above got 7/8 right, while one-word descriptions got 5/8. The one-word version was also confidently wrong: it labelled «اپ باز نمیشه» as cancel with 0.92 confidence.
  • The input can be formal or colloquial Persian, Finglish, or Persian mixed with English words. Arabic ي/ك and missing half-spaces are handled.
  • On a CPU a decision takes roughly 150-200 ms. On a GPU it is much faster.

Full benchmark results: before vs after

Base model: convaiinnovations/laya / multilingual @ 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Scores are from the first of three repeats, and every repeat gave identical predictions. Neither model saw these 64 messages or these two prompts during training.

Persian instructions (choice_fa)

Metric Base This model Change
Accuracy 36/64 51/64 +15
Macro-F1 0.513 0.795
False cancel (would wrongly close an account) 4 2
Missed cancel 9 3
Mean answer confidence 0.84 0.69
ECE (calibration error, lower is better) 0.273 0.103
Family Base This model Change
formal 6/8 8/8 +2
colloquial 5/8 7/8 +2
finglish 4/8 4/8 0
orthography 4/8 6/8 +2
code_mixed 5/8 8/8 +3
negation 3/8 6/8 +3
sarcasm_taarof 5/8 6/8 +1
digits 4/8 6/8 +2
Label (recall) Base This model
billing 13/16 14/16
technical 14/16 14/16
cancel 7/16 13/16
other 2/16 10/16
Confusion matrices (rows = gold, columns = predicted)

Base

billing technical cancel other
billing 13 2 1 0
technical 1 14 1 0
cancel 1 8 7 0
other 6 6 2 2

This model

billing technical cancel other
billing 14 0 1 1
technical 1 14 1 0
cancel 1 2 13 0
other 4 2 0 10

English instructions (choice_en)

Metric Base This model Change
Accuracy 41/64 44/64 +3
Macro-F1 0.633 0.685
False cancel (would wrongly close an account) 2 3
Missed cancel 6 6
Mean answer confidence 0.82 0.71
ECE (calibration error, lower is better) 0.250 0.108
Family Base This model Change
formal 5/8 8/8 +3
colloquial 7/8 6/8 -1
finglish 4/8 5/8 +1
orthography 5/8 5/8 0
code_mixed 6/8 6/8 0
negation 3/8 4/8 +1
sarcasm_taarof 6/8 5/8 -1
digits 5/8 5/8 0
Label (recall) Base This model
billing 10/16 10/16
technical 15/16 14/16
cancel 10/16 10/16
other 6/16 10/16
Confusion matrices (rows = gold, columns = predicted)

Base

billing technical cancel other
billing 10 3 2 1
technical 1 15 0 0
cancel 0 6 10 0
other 4 6 0 6

This model

billing technical cancel other
billing 10 2 2 2
technical 1 14 1 0
cancel 0 4 10 2
other 3 3 0 10

Training

Base checkpoint laya-multilingual (mmBERT-base, 322M parameters) @ 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
Method top mode: decision head + last 4 encoder layers. Token embeddings frozen.
Objective cross-entropy on the gold option, AdamW, fp16 autocast
Data 200 training and 40 dev Persian support messages (50 and 10 per label), all synthetic
Prompts 4 paraphrased questions (2 Persian, 2 English), none of them the benchmark prompts
Robustness options shuffled with probability 0.5; random ي/ك, half-space and digit variants
Model selection best epoch by dev NLL (epoch 6)
Calibration one choice temperature fitted on dev (1.27)
Hardware a single NVIDIA GeForce GTX 1650 (4 GB), about 10 minutes
Software laya 0.3.21, torch 2.14.0+cu126, seed 42

The dev set has 40 messages x 4 training prompts. Dev accuracy went from 0.562 before training to 0.700 at the selected epoch.

Epoch Train loss Dev accuracy Dev NLL
0 (before) - 0.562 1.555
1 1.077 0.581 0.999
2 0.870 0.619 0.917
3 0.800 0.656 0.896
4 0.800 0.663 0.852
5 0.728 0.700 0.839
6 ✅ selected 0.733 0.700 0.810
7 0.700 0.706 0.818
8 0.639 0.706 0.820

A near-duplicate check (character 3-gram Jaccard < 0.5) guarantees that no training or dev message copies a benchmark message.

Limitations

  • The data is small and synthetic. The training set and the benchmark were drafted the same way, with AI assistance. Real customer traffic will be harder, so expect a smaller gain there than the benchmark shows. Evaluate on your own messages before relying on the model.
  • With Persian instructions, these families did not improve: finglish (4/8).
  • With English instructions, these families did not improve: colloquial (6/8), orthography (5/8), code_mixed (6/8), sarcasm_taarof (5/8), digits (5/8).
  • With English instructions the overall change is small (+3 of 64), which is within noise for 64 cases.
  • It is sensitive to how the options are described. Use descriptive option texts (see Tips). Terse one-word options made it much worse, with high confidence.
  • It is specialised to one 4-way decision. Other Laya tasks and languages were not re-evaluated and may have degraded. Use the original Laya checkpoints for those.
  • Confidence was calibrated on 160 dev rows only. Refit the threshold or temperature on your own data before gating decisions on it.

Links

License

Apache-2.0, the same as the base Laya weights.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for tarfandoon/laya-fa-support

Finetuned
(101)
this model

Space using tarfandoon/laya-fa-support 1