Parakh-4B

Parakh-4B

Parakh (परख) means discerning judgement. Parakh-4B is an open decision model for Indian-language text: give it a message and one or more fixed-answer questions (choice, yes/no, score), and it returns a calibrated probability for every answer in one forward pass. It does not generate text.

  • Languages: Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu, Urdu, Assamese, Odia, and romanised Hinglish.
  • Model: Kev-4B (Qwen3.5-4B-Base + LoRA r16 + pointer head, 4.66B parameters) fine-tuned on 13,965 Indian-language records.
  • Author: Yashaswi Galhotra · Licence: CC BY-NC 4.0 (non-commercial)

Results

Headline numbers are on the full official test sets of IndicSentiment and IndicXNLI (both part of the IndicXTREME benchmark; Doddapaneni et al., 2023) and MASSIVE (FitzGerald et al., 2022), which Parakh never trained on. Further tests below cover reasoning (IndicCOPA), paraphrase (IndicXParaphrase), Hinglish (SentiMix), moderation (HateCheck, TextDetox) and support-triage suites.

Against its base model, on 88,589 test questions (sentiment, NLI, intent; 13 languages):

Kev-4B Parakh-4B
accuracy 70.8% 82.4%
calibration error (ECE) 0.199 0.007
wrong despite ≥ 90% confidence 7.1% 1.8%
questions answerable automatically at ≤ 5% error 8.3% 60.5%

Calibration

Against published fine-tuned models, averaged over languages as in the IndicXTREME paper:

task test questions MuRIL IndicBERT v2 (best published) Kev-4B Parakh-4B
IndicSentiment 12,943 85.3 92.8 93.1 94.8
IndicXNLI 55,030 72.9 75.0 68.1 79.9
MASSIVE intent 20,616 77.2 78.8 64.2 81.1
IndicCOPA (task not trained) 8,843 58.5 62.5 74.3 76.3
IndicXParaphrase (task not trained) 20,018 60.8 57.0 87.7 86.5
parameters 237M 278M 4.66B 4.66B

Against published models

These are not like-for-like comparisons. The published encoders were fine-tuned on English data only, one model per task, and are about 15× smaller. Parakh saw in-language training examples for sentiment, NLI and intent (about 90, 300 and 900 per language). The comparison shows where a small fine-tuned open model stands, not that it is a stronger general model.

Hinglish (SentiMix test, 3,000 tweets): 71.7% accuracy, against 68.3% for Kev-4B.

Against hosted decision models: Jev and GPT-6 Luna

Parakh-4B, Jev 1.13 (TypeSafe) and GPT-6 Luna (OpenAI) answered identical questions: a fixed, stratified sample of the official test sets above (100 per task and language), 3,300 COPA, paraphrase and Hinglish questions, and three use-case suites. Accuracy %:

test set questions Kev-4B GPT-6 Luna Jev 1.13 Parakh-4B Parakh − Jev [95% CI]
Sentiment + NLI + intent 5,292 74.2 76.6 80.6 83.6 +3.0 [+1.9, +4.0]
Sentiment 1,300 94.3 97.8 95.3 95.3 +0.0 [−1.2, +1.2]
NLI 3,292 68.4 69.2 75.6 79.8 +4.2 [+2.5, +5.6]
Intent 700 64.0 72.1 76.9 80.0 +3.1 [+0.1, +5.7]
Hinglish sentiment 500 65.6 60.8 62.2 69.2 +7.0 [+2.0, +12.2]
Triage: civic, health, bank 107 87.9 94.4 86.9 90.7 +3.7 [+0.9, +6.6]
COPA reasoning 1,800 75.7 86.7 87.5 75.6 −11.9 [−14.1, −9.9]
Paraphrase 1,000 87.0 91.0 90.4 86.4 −4.0 [−5.6, −2.4]
Moderation (HateCheck + TextDetox) 1,999 67.1 84.6 80.5 70.5 −10.0 [−11.8, −8.0]
All 8,592 shared benchmark questions 8,592 75.5 79.5 82.1 81.4 −0.7 [−1.6, +0.1]

Against hosted decision models

Where Jev and Luna are better: moderation, common-sense reasoning and paraphrase by clear margins, and Luna on sentiment and triage. Across all shared questions Jev and Parakh are level.

Where Parakh is better: NLI, intent and Hinglish, and above all in how far its confidence can be trusted. On the 5,292 trained-task questions, 1.3% of Parakh's ≥ 90%-confidence answers are wrong, against 4.2% for Jev and 15.8% for Luna (calibration error 0.008, 0.064 and 0.186). On Hinglish, 21% of Jev's and 37% of Luna's ≥ 90%-confidence answers are wrong (Parakh: 0.2%).

Confidence you can act on

Open versus hosted. Jev is a closed API. GPT-6 Luna is a proprietary general-purpose model whose size is not disclosed, billed per token, and exposes only its top few log-probabilities. Parakh's weights are open: it runs on your own servers, so messages never leave your infrastructure, it can be fine-tuned further on your own labels, and a given version never changes. Price is not the difference: at October 2026 list prices Jev and Luna cost about 1.5 and 2 US cents per 1,000 questions; one L4 GPU serving Parakh at full load costs about 1.2 cents.

Jev: jev-1.13.0 via TypeSafe's System One API. GPT-6 Luna: OpenAI Chat Completions with reasoning off, probabilities from answer log-probabilities (OpenAI's Decisions API was not available to us). Scored 4–6 October 2026; hosted models change over time. Hosted outputs were used for evaluation only, never as training data. Intervals are paired bootstrap. Parakh was trained on the training splits of sentiment, NLI and intent; the hosted models' training data is not public. Parakh-4B is not affiliated with or endorsed by TypeSafe or OpenAI.

By language

The release announcement is in LAUNCH.md.

Use it for

  • Routing and intent: support tickets, banking and UPI complaints, civic grievances, helpline queries.
  • Sentiment: product and service reviews, social posts, including Hinglish.
  • Grounding checks: does a text support or contradict a claim.
  • Automating the confident cases: accept answers above a confidence threshold and send the rest to a person. Set the threshold on a few hundred of your own labelled messages, since your data will differ from the test sets.

Good to know

  • Parakh is strongest at routing, intent, sentiment and NLI. For moderation, common-sense reasoning and paraphrase, Jev and GPT-6 Luna are stronger (see the comparison above).
  • Real messages are noisier than benchmarks: try it on a sample of your own data and set the confidence threshold there.
  • Keep a person in the loop for decisions about people, such as credit, hiring, legal or medical.

How the results were checked

  • Training, calibration and model selection used only upstream train and dev splits. Every test question was compared with all 47,903 training and calibration records (exact, normalised and near-duplicate matching). The only overlaps were 72 short MASSIVE commands such as "what time is it", which MASSIVE itself repeats across its splits; without them, intent accuracy is 81.0%.
  • Calibration uses one temperature (1.149), fitted on 810 held-out development records.
  • The bf16 server returns the same answer as the fp32 evaluation on 99.8% of benchmark questions.

Running it

git clone https://github.com/jaredpalmer/kev && cd kev
uv run --extra serve python -m kev.serve --run YAlgoG/parakh-4b --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
  "state": "UPI se 2000 bheje, paise kat gaye par receiver ko nahi mile",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this message?",
             "criteria": {"payments": "Failed or wrong transactions", "fraud": "Unauthorised transactions or scams",
                          "cards": "Debit or credit cards", "loans": "Loans and EMIs"}}
  }}'

Uses about 10 GB of GPU memory in bf16.

Training

init jaredpalmer/kev-4b (base Qwen/Qwen3.5-4B-Base)
recipe LoRA r16 / α32 on all linear layers, lr 5e-5, 1 epoch, batch 8 (2 × 4 accumulation), seed 0, plus 2,000 replayed English Kev records
training data records licence
MASSIVE intent, train split 6,300 CC BY 4.0
IndicXNLI, dev split 3,300 CC BY-NC 4.0 (per dataset card)
IndicSentiment, validation split 1,118 CC0
SentiMix (SemEval-2020 Task 9), train split 1,198 OpenRAIL (Hub tag)
Paraphrase and reasoning examples written for this project (templated and AI-assisted) 2,049 original to this project
Kev English replay 2,000 mixed, including non-commercial

Licence and use restrictions

  • Parakh-4B (adapter and head): CC BY-NC 4.0, non-commercial use only, because IndicXNLI and part of Kev's replay data are non-commercial.
  • Following the OpenRAIL terms of the SentiMix data, do not use it for illegal activity, harassment, discrimination, disinformation, surveillance or profiling of individuals, or fully automated decisions that affect people's rights.
  • Base weights: Qwen3.5-4B-Base (Qwen team, Alibaba Cloud) and Kev-4B (Jared Palmer), both Apache-2.0. This repository contains a modified derivative (a LoRA adapter and a retrained head); see NOTICE.md and LICENSE-APACHE-2.0.txt. Not affiliated with or endorsed by either.
  • Test sets were used for evaluation only and are not redistributed.

Acknowledgements

Jared Palmer (Kev), the Qwen team, AI4Bharat (IndicXTREME), Amazon (MASSIVE), the SemEval-2020 Task 9 organisers (SentiMix) and Röttger et al. (Multilingual HateCheck).

Citation

@misc{galhotra2026parakh,
  title  = {Parakh-4B: an open, calibrated decision model for Indian languages},
  author = {Galhotra, Yashaswi},
  year   = {2026},
  url    = {https://huggingface.co/YAlgoG/parakh-4b}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YAlgoG/parakh-4b

Adapter
(77)
this model

Datasets used to train YAlgoG/parakh-4b