Jev-Multilingual - Urdu, Sindhi and Roman Sindhi Decisions

Jev-Multilingual: Urdu, Sindhi & Roman Sindhi Decisions for Python

Jev-Multilingual is a fine-tuned multilingual extension of muhammadnoman76/jev-urdu, extending the Jev decision architecture to Perso-Arabic Sindhi (Ψ³Ω†ΪŒΩŠ) and Roman Sindhi, while preserving the original Urdu decision workflows.

The model is designed for structured decision-making rather than conversational text generation.

Developed by

Upstream Jev-Urdu model and library: Muhammad Noman / LughaatNLP

Sindhi and Roman Sindhi adaptation, presets and evaluation: Shakeel Ahmed Sanjrani

Model Repository | Upstream Model | Apache-2.0 License


Overview

Urdu and Sindhi applications frequently receive text in multiple forms:

  • Perso-Arabic script
  • Romanized text
  • English mixed with local languages
  • informal user-generated text

Jev-Multilingual provides a common typed decision interface for these workflows.

Instead of generating a conversational answer, the model evaluates a message against a structured question and returns a decision together with model probabilities.

For example:

import jev_urdu

model = jev_urdu.load(
    "shakeel143/jev-multilingual"
)

result = model.sentiment(
    "Ω‡ΩŠ سروس ΨͺΩ…Ψ§Ω… Ψ¨Ω‡ΨͺΨ±ΩŠΩ† Ϋ½ سٺي Ψ’Ω‡ΩŠ.",
    lang="sd"
)

print(result)

Roman Sindhi can use the same decision interface:

result = model.sentiment(
    "hee phone daadho sutho aa, battery zabardast aa.",
    lang="sd"
)

print(result)

Urdu remains available through the same API:

result = model.sentiment(
    "یہ ΩΩˆΩ† بہΨͺ Ψ§Ϊ†ΪΎΨ§ ΫΫ’ΨŒ بیٹری بھی Ψ²Ψ¨Ψ±Ψ―Ψ³Ψͺ ہے۔",
    lang="ur"
)

print(result)

The output is a structured decision rather than generated conversational text.


What This Release Adds

1. Sindhi Task Presets

The library adds Sindhi-specific task definitions and labels for workflows including:

  • sentiment
  • triage / priority
  • topic classification
  • claim verification

The prompts and labels are written specifically for Sindhi rather than simply translating the Urdu strings mechanically.


2. Language-Aware API

The extended API allows applications to explicitly select the language:

model.sentiment(text, lang="sd")
model.sentiment(text, lang="ur")

model.triage(text, lang="sd")
model.triage(text, lang="ur")

model.topic(text, lang="sd")

model.check_claim(
    context,
    claim,
    lang="sd"
)

Roman Sindhi is evaluated through the Sindhi language pathway because it represents Sindhi written with Latin characters rather than a separate language.


3. Roman Sindhi Support

A major focus of this adaptation is Roman Sindhi.

Roman Sindhi does not have one universally standardized spelling system. The same expression can therefore appear with different Latin-script spellings.

Examples include:

sutho
sutho aa
daadho sutho
daadho bekaar
mayosi

The adaptation evaluates Roman Sindhi vocabulary and informal spelling patterns rather than treating Sindhi as exclusively Perso-Arabic-script text.


4. Sindhi Contradiction Resolution

The adaptation specifically investigated lexical contradiction and antonym behavior in Sindhi.

One of the observed baseline weaknesses involved the temperature opposition:

Ϊ―Ψ±Ω… ↔ ٿڌي

The post-adaptation evaluation improved the tested contradiction pair from:

0/2 β†’ 2/2

on the controlled evaluation examples.

This result should be interpreted as evidence of improvement on the tested evaluation set, not as proof that the model perfectly understands all Sindhi antonyms.


5. Urdu Regression Protection

The adaptation was evaluated against an Urdu regression suite.

The evaluated Urdu positive, negative and neutral controls remained unchanged in the regression test:

100% preserved on the tested regression suite.

This indicates that no regression was observed on those specific Urdu controls after Sindhi/Roman Sindhi adaptation.


Built-in Tasks

Method Languages Purpose
triage(text, lang="ur"|"sd") Urdu, Sindhi Issue, human-agent request and priority
sentiment(text, lang="ur"|"sd") Urdu, Sindhi, Roman Sindhi Sentiment and dissatisfaction
topic(text, lang="sd") Sindhi, Roman Sindhi Domain/topic classification
check_claim(context, claim, lang="ur"|"sd") Urdu, Sindhi Claim/context relationship
consent(text) Urdu, Sindhi Permission and authorization
detect_injection(text) Urdu, Sindhi Instruction/injection detection

The exact output labels depend on the task preset.


Installation

Recommended: Install the library from this repository

To use the latest Sindhi and Roman Sindhi API implementation:

pip install "git+https://huggingface.co/shakeel143/jev-multilingual#subdirectory=library"

Then:

import jev_urdu

model = jev_urdu.load(
    "shakeel143/jev-multilingual"
)

The package distribution is named:

jev-urdu

while the Python import is:

jev_urdu

Device Selection

GPU is optional.

The loader automatically selects an available device.

You can explicitly choose CPU:

model = jev_urdu.load(
    "shakeel143/jev-multilingual",
    device="cpu"
)

or CUDA:

model = jev_urdu.load(
    "shakeel143/jev-multilingual",
    device="cuda"
)

CUDA requires a compatible NVIDIA GPU and CUDA-enabled PyTorch installation.


Custom Decision Questions

The Jev interface allows applications to define their own structured questions.

For example:

from jev_urdu import choice, yes_no

questions = {
    "topic": choice(
        "Ω‡ΩŠ ΩΎΩŠΨΊΨ§Ω… ΪͺΩ‡Ϊ™ΩŠ Ω…ΩˆΨΆΩˆΨΉ Ψ³Ψ§Ω† Ω„Ψ§Ϊ³Ψ§ΩΎΩŠΩ„ Ψ’Ω‡ΩŠΨŸ",
        [
            "ΨͺΨΉΩ„ΩŠΩ…",
            "ٽيΪͺΩ†Ψ§Ω„Ψ§Ψ¬ΩŠ",
            "Ψ²Ψ±Ψ§ΨΉΨͺ",
            "Ψ΅Ψ­Ψͺ",
            "ٻيو",
        ],
    ),

    "urgent": yes_no(
        "Ϊ‡Ψ§ Ω‡ΩŠ Ω…ΨΉΨ§Ω…Ω„Ωˆ فوري Ψ’Ω‡ΩŠΨŸ"
    ),
}

result = model.ask(
    "Ψ§Ϊ„ ΪͺΪ»Ϊͺ جو Ψ§Ϊ―Ω‡Ω‡ وڌي ويو Ψ’Ω‡ΩŠ.",
    questions,
)

print(result)

This allows developers to create application-specific decision workflows without building a new model head for every classification problem.


Evaluation Methodology

The development process followed a staged evaluation strategy.

Upstream Jev-Urdu
        β”‚
        β–Ό
Zero-shot Sindhi evaluation
        β”‚
        β”œβ”€β”€ Sentiment
        β”œβ”€β”€ Topic
        β”œβ”€β”€ Priority
        β”œβ”€β”€ Claim verification
        └── Lexical contradiction
        β”‚
        β–Ό
Weakness identification
        β”‚
        β–Ό
Sindhi + Roman Sindhi adaptation
        β”‚
        β–Ό
Post-adaptation evaluation
        β”‚
        β”œβ”€β”€ Sindhi improvement
        β”œβ”€β”€ Roman Sindhi evaluation
        └── Urdu regression tests
        β”‚
        β–Ό
Jev-Multilingual

The evaluation suite uses controlled, reproducible examples.

The reported percentages should therefore be interpreted as development benchmarks, not as large-scale language benchmarks.


Baseline vs Adapted Results

Sindhi zero-shot baseline

Before adaptation, the upstream checkpoint produced the following results on the controlled Sindhi evaluation suites:

Evaluation Baseline
Sindhi sentiment 87.5% (21/24)
Sindhi topic classification 71.9% (23/32)
Sindhi priority / triage 75.0% (9/12)
Sindhi claim verification 66.7% (8/12)
Roman Sindhi sentiment 66.7% (8/12)

These results established the baseline against which the adaptation was evaluated.


Adaptation Results

Sindhi lexical contradiction

A controlled contradiction probe involving:

Ϊ―Ψ±Ω… ↔ ٿڌي

improved from:

Before adaptation: 0/2
After adaptation:  2/2

Improvement: 0% β†’ 100% on this controlled pair.


Roman Sindhi negative detection

The Roman Sindhi evaluation included authentic vocabulary such as:

daadho bekaar
mayosi

The tested negative-detection result improved from:

Before adaptation: 50%
After adaptation:  75%

This indicates improvement on the evaluated Roman Sindhi examples.

Because the evaluation set is small, the result should not be interpreted as a general Roman Sindhi benchmark.


Urdu Regression Evaluation

The adapted checkpoint was tested against the Urdu regression controls used during development.

The tested positive, negative and neutral Urdu controls remained correct:

100% β€” 0 observed regression on the evaluated controls.

This provides a regression anchor for the multilingual adaptation.


Important Interpretation of the Results

The reported results are intentionally presented with their evaluation-set sizes.

For example:

2/2

does not mean that the model has solved Sindhi contradiction in general.

It means that the model correctly handled both examples in that particular controlled probe.

Similarly:

75% Roman Sindhi negative detection

does not represent a comprehensive Roman Sindhi benchmark.

Larger independently constructed evaluation datasets are required for stronger claims.


Architecture

Component Specification
Encoder ModernBERT-family bidirectional encoder based on jhu-clsp/mmBERT-base
Encoder layers 22
Hidden size 768
Multilingual vocabulary ~256K tokens
Decision head 2-layer Transformer encoder + type embeddings + marker-level MLP scorer
Parameters ~322M
Input languages/scripts Urdu, Roman Urdu, Sindhi, Roman Sindhi, mixed English
Decision interface Typed classification / decision questions
License Apache-2.0

Adaptation / Fine-Tuning Configuration

The Sindhi/Roman Sindhi adaptation was performed starting from the Jev-Urdu checkpoint.

The training configuration included:

Parameter Value
Optimizer AdamW
Encoder learning rate 2.5e-5
Decision-head learning rate 1e-4
Weight decay 0.01
Token embeddings Frozen
train_embeddings false

The adaptation was designed to improve Sindhi and Roman Sindhi behavior while preserving the upstream Urdu capability.


Model Size

The model contains approximately:

322 million parameters

The checkpoint is stored in Safetensors format.

Actual storage requirements include the checkpoint, tokenizer files, model loading memory and inference tensors.

A GPU is recommended for faster inference but is not required.


Limitations

Small evaluation sets

The reported development benchmarks use controlled datasets and probes.

They should not be interpreted as large-scale language benchmarks.

Roman Sindhi spelling variation

Roman Sindhi has substantial spelling variation. Performance can therefore vary with spelling, morphology and writing style.

Probability calibration

Model probabilities should be validated for the intended application.

A high probability does not guarantee correctness.

Task-specific adaptation

The model has been adapted for the evaluated Sindhi/Roman Sindhi decision workflows. This does not imply that every possible Sindhi NLP task will perform equally well.

Safety-critical decisions

The model should not be used as the sole decision-maker for medical, legal, financial, safety-critical or other consequential decisions.

Application-level validation and human review remain necessary.

Prompt/injection detection

The injection-detection capability is a model behavior and classification feature. It should not be treated as a complete security boundary.


Intended Use

Jev-Multilingual is intended for:

  • Sindhi NLP experimentation
  • Urdu/Sindhi customer-support routing
  • sentiment analysis
  • topic classification
  • structured decision workflows
  • Roman Sindhi experimentation
  • multilingual application prototypes
  • research into low-resource language adaptation

It can be integrated into applications where a structured model decision is more useful than free-form text generation.


Not Intended For

The model should not be treated as:

  • a general conversational assistant
  • a translation system
  • a replacement for human decision-making
  • a safety/security guarantee
  • a medical diagnostic system
  • a legal decision system
  • a fully comprehensive Sindhi language understanding benchmark

Reproducibility

The project maintains its development workflow through Git and Hugging Face.

The repository contains:

  • model checkpoint
  • tokenizer
  • Python library
  • tests
  • Sindhi evaluation code
  • adaptation-related code
  • documentation

Evaluation notebooks are used to reproduce baseline and post-adaptation experiments.


Future Work

Planned directions include:

  1. Larger Sindhi evaluation datasets.
  2. Larger Roman Sindhi evaluation datasets.
  3. More systematic spelling normalization.
  4. More Sindhi task presets.
  5. Expanded code-mixed Sindhi/English evaluation.
  6. Calibration analysis on Sindhi.
  7. Cross-domain evaluation.
  8. Independent held-out benchmarks.
  9. Additional Sindhi lexical and semantic evaluations.
  10. Upstream contribution of the reusable Sindhi extensions where appropriate.

Lineage

jhu-clsp/mmBERT-base
        β”‚
        β–Ό
muhammadnoman76/jev-urdu
        β”‚
        β–Ό
shakeel143/jev-multilingual
        β”‚
        β”œβ”€β”€ Sindhi adaptation
        β”œβ”€β”€ Roman Sindhi adaptation
        β”œβ”€β”€ Sindhi task presets
        β”œβ”€β”€ Evaluation suite
        └── Urdu regression tests

The upstream Jev-Urdu model and library were developed by Muhammad Noman through LughaatNLP.

The underlying mmBERT encoder was developed by JHU CLSP.

Sindhi/Roman Sindhi adaptation, presets, evaluation and multilingual extension work were developed by Shakeel Ahmed Sanjrani.


Relationship to TypeSafe Jev

This project is based on the open Jev-Urdu project and is not affiliated with TypeSafe's commercial Jev API.


License and Attribution

This project uses the Apache-2.0 license, subject to the licenses and attribution requirements of the upstream components.

See:

library/LICENSE
library/NOTICE

Please retain the appropriate upstream attribution when redistributing the library or model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for shakeel143/jev-multilingual

Finetuned
(1)
this model