Instructions to use YAlgoG/parakh-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use YAlgoG/parakh-4b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-4B-Base") model = PeftModel.from_pretrained(base_model, "YAlgoG/parakh-4b") - Notebooks
- Google Colab
- Kaggle
Parakh-4B
Parakh (परख) means discerning judgement. Parakh-4B is an open decision model for Indian-language text: give it a message and one or more fixed-answer questions (choice, yes/no, score), and it returns a calibrated probability for every answer in one forward pass. It does not generate text.
- Languages: Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu, Urdu, Assamese, Odia, and romanised Hinglish.
- Model: Kev-4B (Qwen3.5-4B-Base + LoRA r16 + pointer head, 4.66B parameters) fine-tuned on 13,965 Indian-language records.
- Author: Yashaswi Galhotra · Licence: CC BY-NC 4.0 (non-commercial)
Results
Headline numbers are on the full official test sets of IndicSentiment and IndicXNLI (both part of the IndicXTREME benchmark; Doddapaneni et al., 2023) and MASSIVE (FitzGerald et al., 2022), which Parakh never trained on. Further tests below cover reasoning (IndicCOPA), paraphrase (IndicXParaphrase), Hinglish (SentiMix), moderation (HateCheck, TextDetox) and support-triage suites.
Against its base model, on 88,589 test questions (sentiment, NLI, intent; 13 languages):
| Kev-4B | Parakh-4B | |
|---|---|---|
| accuracy | 70.8% | 82.4% |
| calibration error (ECE) | 0.199 | 0.007 |
| wrong despite ≥ 90% confidence | 7.1% | 1.8% |
| questions answerable automatically at ≤ 5% error | 8.3% | 60.5% |
Against published fine-tuned models, averaged over languages as in the IndicXTREME paper:
| task | test questions | MuRIL | IndicBERT v2 (best published) | Kev-4B | Parakh-4B |
|---|---|---|---|---|---|
| IndicSentiment | 12,943 | 85.3 | 92.8 | 93.1 | 94.8 |
| IndicXNLI | 55,030 | 72.9 | 75.0 | 68.1 | 79.9 |
| MASSIVE intent | 20,616 | 77.2 | 78.8 | 64.2 | 81.1 |
| IndicCOPA (task not trained) | 8,843 | 58.5 | 62.5 | 74.3 | 76.3 |
| IndicXParaphrase (task not trained) | 20,018 | 60.8 | 57.0 | 87.7 | 86.5 |
| parameters | 237M | 278M | 4.66B | 4.66B |
These are not like-for-like comparisons. The published encoders were fine-tuned on English data only, one model per task, and are about 15× smaller. Parakh saw in-language training examples for sentiment, NLI and intent (about 90, 300 and 900 per language). The comparison shows where a small fine-tuned open model stands, not that it is a stronger general model.
Hinglish (SentiMix test, 3,000 tweets): 71.7% accuracy, against 68.3% for Kev-4B.
Against hosted decision models: Jev and GPT-6 Luna
Parakh-4B, Jev 1.13 (TypeSafe) and GPT-6 Luna (OpenAI) answered identical questions: a fixed, stratified sample of the official test sets above (100 per task and language), 3,300 COPA, paraphrase and Hinglish questions, and three use-case suites. Accuracy %:
| test set | questions | Kev-4B | GPT-6 Luna | Jev 1.13 | Parakh-4B | Parakh − Jev [95% CI] |
|---|---|---|---|---|---|---|
| Sentiment + NLI + intent | 5,292 | 74.2 | 76.6 | 80.6 | 83.6 | +3.0 [+1.9, +4.0] |
| Sentiment | 1,300 | 94.3 | 97.8 | 95.3 | 95.3 | +0.0 [−1.2, +1.2] |
| NLI | 3,292 | 68.4 | 69.2 | 75.6 | 79.8 | +4.2 [+2.5, +5.6] |
| Intent | 700 | 64.0 | 72.1 | 76.9 | 80.0 | +3.1 [+0.1, +5.7] |
| Hinglish sentiment | 500 | 65.6 | 60.8 | 62.2 | 69.2 | +7.0 [+2.0, +12.2] |
| Triage: civic, health, bank | 107 | 87.9 | 94.4 | 86.9 | 90.7 | +3.7 [+0.9, +6.6] |
| COPA reasoning | 1,800 | 75.7 | 86.7 | 87.5 | 75.6 | −11.9 [−14.1, −9.9] |
| Paraphrase | 1,000 | 87.0 | 91.0 | 90.4 | 86.4 | −4.0 [−5.6, −2.4] |
| Moderation (HateCheck + TextDetox) | 1,999 | 67.1 | 84.6 | 80.5 | 70.5 | −10.0 [−11.8, −8.0] |
| All 8,592 shared benchmark questions | 8,592 | 75.5 | 79.5 | 82.1 | 81.4 | −0.7 [−1.6, +0.1] |
Where Jev and Luna are better: moderation, common-sense reasoning and paraphrase by clear margins, and Luna on sentiment and triage. Across all shared questions Jev and Parakh are level.
Where Parakh is better: NLI, intent and Hinglish, and above all in how far its confidence can be trusted. On the 5,292 trained-task questions, 1.3% of Parakh's ≥ 90%-confidence answers are wrong, against 4.2% for Jev and 15.8% for Luna (calibration error 0.008, 0.064 and 0.186). On Hinglish, 21% of Jev's and 37% of Luna's ≥ 90%-confidence answers are wrong (Parakh: 0.2%).
Open versus hosted. Jev is a closed API. GPT-6 Luna is a proprietary general-purpose model whose size is not disclosed, billed per token, and exposes only its top few log-probabilities. Parakh's weights are open: it runs on your own servers, so messages never leave your infrastructure, it can be fine-tuned further on your own labels, and a given version never changes. Price is not the difference: at October 2026 list prices Jev and Luna cost about 1.5 and 2 US cents per 1,000 questions; one L4 GPU serving Parakh at full load costs about 1.2 cents.
Jev: jev-1.13.0 via TypeSafe's System One API. GPT-6 Luna: OpenAI Chat Completions with reasoning off,
probabilities from answer log-probabilities (OpenAI's Decisions API was not available to us). Scored 4–6 October 2026;
hosted models change over time. Hosted outputs were used for evaluation only, never as training data. Intervals are
paired bootstrap. Parakh was trained on the training splits of sentiment, NLI and intent; the hosted models' training
data is not public. Parakh-4B is not affiliated with or endorsed by TypeSafe or OpenAI.
The release announcement is in LAUNCH.md.
Use it for
- Routing and intent: support tickets, banking and UPI complaints, civic grievances, helpline queries.
- Sentiment: product and service reviews, social posts, including Hinglish.
- Grounding checks: does a text support or contradict a claim.
- Automating the confident cases: accept answers above a confidence threshold and send the rest to a person. Set the threshold on a few hundred of your own labelled messages, since your data will differ from the test sets.
Good to know
- Parakh is strongest at routing, intent, sentiment and NLI. For moderation, common-sense reasoning and paraphrase, Jev and GPT-6 Luna are stronger (see the comparison above).
- Real messages are noisier than benchmarks: try it on a sample of your own data and set the confidence threshold there.
- Keep a person in the loop for decisions about people, such as credit, hiring, legal or medical.
How the results were checked
- Training, calibration and model selection used only upstream train and dev splits. Every test question was compared with all 47,903 training and calibration records (exact, normalised and near-duplicate matching). The only overlaps were 72 short MASSIVE commands such as "what time is it", which MASSIVE itself repeats across its splits; without them, intent accuracy is 81.0%.
- Calibration uses one temperature (1.149), fitted on 810 held-out development records.
- The bf16 server returns the same answer as the fp32 evaluation on 99.8% of benchmark questions.
Running it
git clone https://github.com/jaredpalmer/kev && cd kev
uv run --extra serve python -m kev.serve --run YAlgoG/parakh-4b --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "UPI se 2000 bheje, paise kat gaye par receiver ko nahi mile",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this message?",
"criteria": {"payments": "Failed or wrong transactions", "fraud": "Unauthorised transactions or scams",
"cards": "Debit or credit cards", "loans": "Loans and EMIs"}}
}}'
Uses about 10 GB of GPU memory in bf16.
Training
| init | jaredpalmer/kev-4b (base Qwen/Qwen3.5-4B-Base) |
| recipe | LoRA r16 / α32 on all linear layers, lr 5e-5, 1 epoch, batch 8 (2 × 4 accumulation), seed 0, plus 2,000 replayed English Kev records |
| training data | records | licence |
|---|---|---|
| MASSIVE intent, train split | 6,300 | CC BY 4.0 |
| IndicXNLI, dev split | 3,300 | CC BY-NC 4.0 (per dataset card) |
| IndicSentiment, validation split | 1,118 | CC0 |
| SentiMix (SemEval-2020 Task 9), train split | 1,198 | OpenRAIL (Hub tag) |
| Paraphrase and reasoning examples written for this project (templated and AI-assisted) | 2,049 | original to this project |
| Kev English replay | 2,000 | mixed, including non-commercial |
Licence and use restrictions
- Parakh-4B (adapter and head): CC BY-NC 4.0, non-commercial use only, because IndicXNLI and part of Kev's replay data are non-commercial.
- Following the OpenRAIL terms of the SentiMix data, do not use it for illegal activity, harassment, discrimination, disinformation, surveillance or profiling of individuals, or fully automated decisions that affect people's rights.
- Base weights: Qwen3.5-4B-Base (Qwen team, Alibaba Cloud) and Kev-4B (Jared Palmer), both Apache-2.0. This repository
contains a modified derivative (a LoRA adapter and a retrained head); see
NOTICE.mdandLICENSE-APACHE-2.0.txt. Not affiliated with or endorsed by either. - Test sets were used for evaluation only and are not redistributed.
Acknowledgements
Jared Palmer (Kev), the Qwen team, AI4Bharat (IndicXTREME), Amazon (MASSIVE), the SemEval-2020 Task 9 organisers (SentiMix) and Röttger et al. (Multilingual HateCheck).
Citation
@misc{galhotra2026parakh,
title = {Parakh-4B: an open, calibrated decision model for Indian languages},
author = {Galhotra, Yashaswi},
year = {2026},
url = {https://huggingface.co/YAlgoG/parakh-4b}
}
- Downloads last month
- -
Model tree for YAlgoG/parakh-4b
Base model
Qwen/Qwen3.5-4B-Base




