Instructions to use developerJenis/Artha-v1-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use developerJenis/Artha-v1-E2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="developerJenis/Artha-v1-E2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("developerJenis/Artha-v1-E2B") model = AutoModelForMultimodalLM.from_pretrained("developerJenis/Artha-v1-E2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use developerJenis/Artha-v1-E2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "developerJenis/Artha-v1-E2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "developerJenis/Artha-v1-E2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/developerJenis/Artha-v1-E2B
- SGLang
How to use developerJenis/Artha-v1-E2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "developerJenis/Artha-v1-E2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "developerJenis/Artha-v1-E2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "developerJenis/Artha-v1-E2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "developerJenis/Artha-v1-E2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use developerJenis/Artha-v1-E2B with Docker Model Runner:
docker model run hf.co/developerJenis/Artha-v1-E2B

Artha v1 E2B
Artha writes the way people in India actually write. Give it a rough request in English, Hinglish or Gujlish, and it returns the finished email, WhatsApp message or LinkedIn post, in a voice that reads like a real person sent it, not an AI.
Artha v1 E2B is the compact, on-device member of the family. It runs fully offline on a phone through Google AI Edge Gallery, and on any laptop through Ollama or llama.cpp.
| Repository | Format | Use it for |
|---|---|---|
| Artha-v1-E2B | Safetensors (BF16) | Transformers, further fine-tuning |
| Artha-v1-E2B-GGUF | GGUF Q4_K_M | Ollama, llama.cpp, LM Studio |
| Artha-v1-E2B-LiteRT | .litertlm (4-bit, 2.8 GB) |
Google AI Edge Gallery on Android and iOS |
What it does
- Understands how people type. Clean English, rushed one-liners full of typos, Hinglish, Gujlish.
- Returns the message itself, ready to send. Emails come with a subject line and sign-off. WhatsApp replies stay short. Advice questions get a direct, practical answer.
- Sounds human. Gets to the point, uses your specifics, and avoids the usual AI tells: "I hope this email finds you well", "I wanted to reach out", "Furthermore", heavy bullet points.
- Knows Indian workplace English. Lakh and crore, "kindly", "PFA", "EOD", "Sir/Ma'am", "ji", used where they fit and never overdone.
- Tone control.
formal-indian,ceo-crisp,firm-polite,warm-personal,casual-whatsapp.
Examples
Real outputs from Artha v1 E2B on requests it never saw in training.
Email, firm follow-up (on-device, 4-bit build)
Request: client ne 45 din se 2.4 lakh ka payment nahi kiya, firm mail likho. client Mehta Traders, contact Rakesh, mera naam Jenis
Subject: Overdue payment of Rs 2.4 lakh
Hi Rakesh,
This is regarding the invoice for Rs 2.4 lakh which is now 45 days overdue.
We have not received the payment despite our previous reminders. Kindly clear this by EOD today.
Do let me know once the transfer is done so we can close the file.
Jenis
WhatsApp, casual
Request: Pls tell me how to text Rohan that I can't make it to his party in Bandra tonight bcoz I stuck in Mumbai traffic and my Uber cost 450 rs.
Rohan, sorry can't make it to your party in Bandra tonight. Stuck in Mumbai traffic and Uber just cost 450.
Will catch you soon, maybe next weekend?
Advice, Hinglish
Request: interview me 5 baar reject ho gaya, bahut low feel ho raha hai
5 baar rejection sunne mein dard hota hai, aur yeh normal hai ki ab low feel ho raha ho.
Par ek baat yaad rakhna, har rejection ek opportunity hota hai. Ho sakta hai interview ka time sahi na ho, ya woh company ki need alag ho. Yeh tumhari capability ke baare mein nahi batata.
Abhi thoda pause le lo, khud ko recover karne do. Jab tum ready ho, toh sirf 2 ya 3 interviews lo. Is baar thoda confidence se jao, kyunki tumne pehle bhi face kiya hai.
Aur haan, tum akele nahi ho. Yeh journey tumhari value kam nahi karta.
Quick start
On your phone (offline)
- Install Google AI Edge Gallery from the Play Store or App Store (Android 12+, iOS 17+).
- In the model list, choose import from Hugging Face and enter
developerJenis/Artha-v1-E2B-LiteRT. - Pick the GPU backend. After the one-time 2.8 GB download, everything runs on the device.
On a laptop with Ollama
ollama run hf.co/developerJenis/Artha-v1-E2B-GGUF:Q4_K_M
With llama.cpp
llama-cli -hf developerJenis/Artha-v1-E2B-GGUF:Q4_K_M --jinja
With Transformers (5.5 or newer)
import torch
from transformers import AutoProcessor, AutoModelForMultimodalLM
model_id = "developerJenis/Artha-v1-E2B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": [{"type": "text", "text":
"client ne 45 din se 2.4 lakh ka payment nahi kiya, firm mail likho. client Mehta Traders, contact Rakesh, mera naam Jenis"}]}]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=400, do_sample=True, temperature=0.7, top_p=0.9)
print(processor.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Prompting
A plain request works. Put in the details you want used: names, amounts, dates, what already happened.
For explicit control over channel and tone, add this system prompt. It is the exact format used in training (v1 was trained under its development name, Desi Draft; the name has no effect on output).
You are Desi Draft, a writing assistant for India. Write the way a real Indian professional or friend would, not like an AI.
Channel: email
Tone: ceo-crisp
| Setting | Options |
|---|---|
Channel |
email, whatsapp, linkedin, chat |
Tone |
formal-indian, ceo-crisp, firm-polite, warm-personal, casual-whatsapp |
Recommended sampling: temperature 0.7, top_p 0.9, up to 400 new tokens.
Evaluation
On 40 held-out requests, outputs were scored with an automatic human-style checker that penalises AI phrasing, dashes, markdown, assistant preambles, placeholder brackets, overlong messages and uniform sentence rhythm (0 to 100, higher reads more human).
| Model | Human-style score |
|---|---|
| Gemma 4 E2B (base) | 52.3 |
| Artha v1 E2B | 99.0 |
This is an automatic style metric, not a human judgement. It measures how much the writing avoids AI patterns, not whether every fact is right. A blind human preference test is planned for the next release.
Training
| Base model | google/gemma-4-E2B-it, loaded through Unsloth |
| Stage 1: SFT | 4,227 examples, 2 epochs, LoRA rank 32 (16-bit), loss on responses only |
| Stage 2: DPO | 4,087 preference pairs, 1 epoch, beta 0.1 |
| Hardware | 1x NVIDIA A100 40 GB |
Data. Requests cover 50 everyday Indian scenarios (payment follow-ups, leave requests, salary negotiation, resignations, team memos, investor updates, bank escalations, condolences, festival wishes, career advice) across four channels and six typing styles. Human-style answers were written by an open teacher model, Qwen 3.5 27B, guided by 59 hand-written reference examples. For DPO, each human-style answer was paired with a generic AI-assistant answer to the same request. Answers were filtered for AI phrasing, language mismatch, missing subject lines and advice given in place of a draft.
Limitations
- Check the details before sending. v1 sometimes adds specifics you did not give it, such as a deadline or a previous reminder (see the first example above). This is the main focus of the next release.
- Synthetic training data. Style reflects urban professional Indian English and common Hinglish and Gujlish; other regional varieties are less covered.
- Not an expert system. No knowledge of current events. Do not rely on it for legal, medical or financial decisions.
- Text-focused. The base model's image and audio abilities are inherited but were not tuned.
License
Apache 2.0, the same license as the Gemma 4 base model.
Author
Built by Jenis (developerJenis).
- Downloads last month
- 13
Model tree for developerJenis/Artha-v1-E2B
Base model
google/gemma-4-E2B