Instructions to use ajayk0608/parchi-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ajayk0608/parchi-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ajayk0608/parchi-v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ajayk0608/parchi-v2") model = AutoModelForCausalLM.from_pretrained("ajayk0608/parchi-v2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ajayk0608/parchi-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ajayk0608/parchi-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ajayk0608/parchi-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ajayk0608/parchi-v2
- SGLang
How to use ajayk0608/parchi-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ajayk0608/parchi-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ajayk0608/parchi-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ajayk0608/parchi-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ajayk0608/parchi-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ajayk0608/parchi-v2 with Docker Model Runner:
docker model run hf.co/ajayk0608/parchi-v2
parchi-v2
Parchi reads the purchase note a grocery (kirana) shopkeeper types to a supplier, in Hinglish, English or Telugu (script or romanised), and returns the order as JSON: vendor, plus item, quantity, unit and rupee rate for each line.
It is Qwen2.5-0.5B-Instruct fine-tuned with LoRA (r=16) on 2,600 generated notes over a 64-item catalogue, trained on a laptop CPU, then merged. Scores are exact match on notes written by hand that the model never saw: every field of every line has to be right. Hinglish set: 34 notes; Telugu set: 36 notes.
This repo is the full model (16-bit safetensors) for Python / transformers. For phones, laptops, llama.cpp, Ollama or LM Studio use the quantized builds in ajayk0608/parchi-v2-GGUF.
Which build should I use?
| File | Build | Size | Best for | Hinglish exact | Telugu exact |
|---|---|---|---|---|---|
| this repo | Full model (safetensors) | 953 MB | Developers in Python, servers, and anyone fine-tuning further | 74% | 64% |
| parchi-v2-LoRA | LoRA adapter (PEFT) | 44 MB | Builders who fine-tune further or already have Qwen2.5-0.5B-Instruct | 74% | 64% |
parchi-v2-f16.gguf |
F16 GGUF | 948 MB | People who want to make their own quantized builds | — | — |
parchi-v2-q8_0.gguf |
Q8_0 | 506 MB | Laptops and desktops, shop-counter PCs, servers without a GPU | 74% | 64% |
parchi-v2-q6_k.gguf |
Q6_K | 482 MB | Newer phones (6 GB RAM or more) and small office PCs | 74% | 67% |
parchi-v2-q5_k_m.gguf |
Q5_K_M | 400 MB | Mid-range phones (4 to 6 GB RAM) that want accuracy first | 82% | 69% |
parchi-v2-q4_k_m.gguf ⭐ |
Q4_K_M | 379 MB | Most phones: the default choice for an on-device app | 76% | 75% |
parchi-v2-q4_0.gguf |
Q4_0 | 335 MB | Older or budget Android phones (3 to 4 GB RAM), Raspberry Pi | 68% | 42% |
⭐ = the default for a phone app.
The hand-written sets are small (34 and 36 notes), so one note moves a score by about 3 points. Builds within a few points of each other perform the same, and a quantized build scoring above the full model is noise, not an improvement.
Prompt
Always send this system prompt (the model was trained with it), then the note as the user message. Use greedy decoding (temperature 0).
You read a kirana (grocery) shop purchase note, often in Hinglish, and return the order as JSON only. Schema: {"vendor": string or null, "lines": [{"item": string, "qty": number, "unit": string, "rate": number or null}]}. item is the canonical English name; unit is one of kg, g, l, ml, pcs, dozen, packet, bottle, bunch, box; rate is the rupee price per unit if stated, else null; vendor is who the note is addressed to, else null.
Example output:
{"vendor":"Sharma Store","lines":[{"item":"wheat flour","qty":10,"unit":"kg","rate":null},{"item":"mustard oil","qty":1,"unit":"l","rate":165}]}
Use with transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("ajayk0608/parchi-v2")
model = AutoModelForCausalLM.from_pretrained("ajayk0608/parchi-v2")
SYSTEM = "..." # the system prompt above
msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": "bhai 5 kg chini aur 2 dzn ande bhej do, chini 42 ka"}]
enc = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**enc, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0, enc["input_ids"].shape[1]:], skip_special_tokens=True))
Limits
- Knows the 64 catalogue items and their spellings; an unknown item is guessed onto a known one.
- One order per note, up to about six lines; rates are per unit as written.
- Trained on generated notes: real notes can surface patterns it has not seen (common misses: the next line's number taken as a rate, near-neighbour Telugu spellings, brand-only items).
- Downloads last month
- 605