Instructions to use mohameddalii/coda-llm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mohameddalii/coda-llm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mohameddalii/coda-llm") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("mohameddalii/coda-llm") model = AutoModelForCausalLM.from_pretrained("mohameddalii/coda-llm", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mohameddalii/coda-llm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mohameddalii/coda-llm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohameddalii/coda-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mohameddalii/coda-llm
- SGLang
How to use mohameddalii/coda-llm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mohameddalii/coda-llm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohameddalii/coda-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mohameddalii/coda-llm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohameddalii/coda-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mohameddalii/coda-llm with Docker Model Runner:
docker model run hf.co/mohameddalii/coda-llm
Coda LLM (Granite-4.2-8B Najdi Sales Agent)
Coda LLM is an enterprise-grade Saudi sales and voice-agent model built on top of ibm-granite/granite-4.2-8b. It is specifically fine-tuned for high-velocity real-time phone calls, tele-sales, and customer relationship interactions in Saudi Arabia.
Key Capabilities & Pillars
Native Najdi Dialect with Full Tashkeel (TTS-Grade):
- Natural Saudi business and everyday Najdi vernacular ("أبشر", "يا هلا والله", "وش الباقة اللي تناسبك؟", "تمام كذا").
- Diacritized outputs optimized directly for TTS (Text-to-Speech) engines with clean phonetic clarity and zero MSA robotic leaks.
Native Function Calling & Tool Use:
- Evaluated across real-world enterprise sales tools (
search_products,calculate_quote,schedule_appointment,create_subscription,check_order_status). - 100% negative rejection rate on irrelevant queries (e.g. general chit-chat, weather) — never hallucinates unrequested tool calls.
- 100% grounded pricing: strictly extracts price quotes from tool outputs instead of inventing figures.
- Evaluated across real-world enterprise sales tools (
Autonomous Sales Intelligence:
- Discovery & Qualification: Actively asks clarifying questions about needs, budget, and family/team size before proposing packages.
- Objection Handling: Handles price objections via value framing and flexible plans instead of discount dumping.
- Immediate Stop-Selling: Strictly respects customer rejection cues ("ماني مهتم", "لا تتصلون علي") with 100% immediate graceful exit.
Reasoning & Deliberation:
- Generates internal thinking thoughts (
<think>...</think>) before tool calling or conversational delivery to ensure multi-step sales constraints are honored.
- Generates internal thinking thoughts (
Benchmark Results (v10 Production Release)
Evaluated across a comprehensive 4-pillar, 27-scenario multi-trial suite:
| Pillar | Test Category | Score | Details |
|---|---|---|---|
| Tool Calling | Core Sales Tools (search_products, create_subscription, check_order_status) |
100% | Flawless JSON argument generation |
| Negative Rejection | Non-tool queries (Chit-chat, Food, Weather) | 100% | Zero false-positive tool calls |
| Reasoning | Multi-constraint arithmetic & logic puzzles | 88% | 0.0% truncation rate (converges cleanly within token limits) |
| Sales Discipline | Immediate Stop-Selling (sl_stop_001, sl_stop_002) |
95% - 100% | Aborts sales immediately upon customer request |
| Sales Discipline | Value Framing vs Discount Dumping | 100% | Protects product value |
| Sales Discipline | Retention & Objection Handling | 100% | Resolves hesitation without aggressive pressure |
Quickstart & Usage
1. Serving with vLLM (Recommended for Production & Voice Pipelines)
Run the OpenAI-compatible vLLM server:
python -m vllm.entrypoints.openai.api_server \
--model mohameddalii/coda-llm \
--served-model-name coda-llm \
--tool-call-parser qwen3_coder \
--reasoning-parser deepseek_r1 \
--enable-auto-tool-choice \
--port 8000
Note on Speculative Decoding: Empirical load-testing demonstrates that n-gram speculative decoding offers no speedup on diacritized Arabic text due to novel morphological generation. Serve in standard mode for optimal throughput and low latency.
2. Multi-turn Python Client (with Function Calling)
import json
import requests
API_URL = "http://localhost:8000/v1/chat/completions"
tools = [
{
"type": "function",
"function": {
"name": "search_products",
"description": "Search available telecom and internet packages.",
"parameters": {
"type": "object",
"properties": {
"category": {"type": "string", "description": "e.g. fiber, 5g, business"}
},
"required": ["category"]
}
}
},
{
"type": "function",
"function": {
"name": "calculate_quote",
"description": "Calculate exact subscription cost.",
"parameters": {
"type": "object",
"properties": {
"product_id": {"type": "string"},
"quantity": {"type": "integer"}
},
"required": ["product_id"]
}
}
}
]
messages = [
{"role": "system", "content": "أنت ممثل مبيعات اتصالات بلهجة نجدية طبيعية مع التشكيل الكامل."},
{"role": "user", "content": "وش عندكم عروض فايبر للمنزل؟"}
]
payload = {
"model": "coda-llm",
"messages": messages,
"tools": tools,
"temperature": 0.7,
"max_tokens": 1024
}
response = requests.post(API_URL, json=payload).json()
assistant_message = response["choices"][0]["message"]
print("Assistant Response:", assistant_message)
3. Loading with Hugging Face Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "mohameddalii/coda-llm"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
messages = [
{"role": "system", "content": "أنت ممثل مبيعات سعودي بلهجة نجدية طبيعية مع التشكيل الكامل."},
{"role": "user", "content": "السلام عليكم، وش أرخص باقة عندكم؟"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Training Details
- Base Model:
ibm-granite/granite-4.2-8b - Method: QLoRA Fine-Tuning merged directly into full 16-bit safetensors weights.
- Dataset: Highly curated dataset of real multi-turn Saudi tele-sales conversations, anti-hallucination negative examples, tool calling schemas, and diacritized speech synthesis pairs.
- Context Length: Supports up to 131,072 tokens natively.
- Downloads last month
- 241
Model tree for mohameddalii/coda-llm
Base model
ibm-granite/granite-4.1-8b-base