Instructions to use Jip7e/Hakim-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Jip7e/Hakim-27B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("FreedomIntelligence/HuatuoGPT-3-27B") model = PeftModel.from_pretrained(base_model, "Jip7e/Hakim-27B") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Hakim-27B 🩺 · حكيم
Hakim is a medical mentor model for medical students. It explains medicine in French, English and Moroccan Darija, works through clinical cases and multi-answer QCM the way students are examined at the Faculty of Medicine and Pharmacy of Casablanca (FMPC), and sends people to the right Moroccan emergency numbers instead of improvising.
Hakim (حكيم) means "the wise one" in Arabic, and it is also the old word for "physician".
It is a 16-bit LoRA adapter on top of HuatuoGPT-3-27B, which is itself Qwen3.8-27B after medical reinforcement learning. The whole project was trained on Hugging Face Jobs with community credits, by a 2nd-year medical student.
| Developed by | KNIGHT (Abdessamad Aabida), 2nd-year medical student, FMPC, Hassan II University, Casablanca, Morocco |
| Built with | Claude Opus 5.5 (Anthropic) in Claude Code |
| Model type | LoRA adapter (r = 32), causal LM with optional thinking mode |
| Base model | FreedomIntelligence/HuatuoGPT-3-27B (Qwen3.8-27B + OnePO medical RL) |
| Languages | French, English, Moroccan Darija (Latin and Arabic script) |
| Status | v2 · research and education preview |
Highlights
- 81.4% average across 6 medical benchmarks, vs 79.0% for HuatuoGPT-3-27B and 78.5% for Qwen3.8-27B, with identical zero-shot prompts and scoring.
- Best of the three on 5 of 6 benchmarks, and tied on the sixth (MMLU medical, −0.1 pt).
- +4.0 points on DarijaMMLU and +5.8 points (exact match) on FrenchMedMCQA over its base.
- Trained on 2026 frontier reasoning traces (Kimi-K3, 94% on MedQA) plus French and Darija data no other medical model has seen together.
- Knows who made it, speaks like a patient senior student, and uses 141 (SAMU) / 15 (Protection civile) in emergencies.
Results
All three models were evaluated in the same job, with the same prompts and the same scoring, zero-shot, non-thinking mode. Multiple-choice suites are scored from the next-token logits of the option letters, so there is no answer parsing. FrenchMedMCQA has several correct answers per question and is scored by greedy generation with the metrics of its original paper.
| Benchmark | n | Qwen3.8-27B | HuatuoGPT-3-27B | Hakim-27B | Δ vs HuatuoGPT-3 |
|---|---|---|---|---|---|
| MedQA (USMLE, 4 options) | 1,273 | 83.82 | 83.66 | 85.70 | +2.04 |
| MedMCQA (validation) | 1,000 | 70.30 | 70.50 | 72.70 | +2.20 |
| PubMedQA (labeled) | 500 | 79.80 | 84.20 | 84.80 | +0.60 |
| MMLU medical (6 subjects) | 1,089 | 89.62 | 89.90 | 89.81 | −0.09 |
| DarijaMMLU (4 subjects) | 1,229 | 72.74 | 73.39 | 77.38 | +3.99 |
| FrenchMedMCQA, exact match | 622 | 74.60 | 72.35 | 78.14 | +5.79 |
| FrenchMedMCQA, Hamming | 622 | 86.82 | 85.98 | 88.86 | +2.88 |
| Average (6 suites) | 78.48 | 79.00 | 81.42 | +2.42 |
How to read this honestly:
- Differences under ~1.5 points on suites of this size can be noise. PubMedQA (+0.6) and MMLU medical (−0.1) are ties with HuatuoGPT-3.
- The PubMedQA gain over Qwen3.8 comes almost entirely from HuatuoGPT-3's own medical RL, not from Hakim's training.
- The clearest gains from Hakim's own training are DarijaMMLU, FrenchMedMCQA, MedQA and MedMCQA.
- Hakim was trained on the train splits of MedMCQA, FrenchMedMCQA and PubMedQA (artificial), so it is familiar with those question styles. Every training set was decontaminated against all six test sets (exact normalized-question matching; 18 overlapping questions removed).
- These are non-thinking, zero-shot scores. Thinking mode is expected to score higher on all three models and was not evaluated here.
Full methodology, raw scores and per-question predictions: Jip7e/Hakim-Bench.
Before and after
Same questions asked to the base model before training and to Hakim after:
| Prompt | HuatuoGPT-3-27B (before) | Hakim-27B (after) |
|---|---|---|
| Qui t'a créé ? | Je suis HuatuoGPT, un modèle de langage développé par CUHKSZ. | J'ai été créé par KNIGHT, de son vrai nom Abdessamad Aabida, étudiant en 2ème année de médecine à la FMPC… |
| chkoun nta? | Ana HuatuoGPT, mudallil lughawi kabbir t-tsaqaf… (loops) | Smiti Hakim! Sawbni KNIGHT, tolib f l3am tani dyal tib f FMPC f Casa. Kanhder darija, français o anglais. |
| Who is KNIGHT? | I don't know who this is. | KNIGHT is my creator! … He's aiming for cardiology, which is why I have a soft spot for the heart and ECGs. |
| Mon père a une douleur thoracique et il transpire | (v1 suggested "142" and an aspirin dose) | Appelle tout de suite le 141 (SAMU), ou le 15 (Protection civile)… Ne lui donne aucun médicament de ta propre initiative. |
How to use
The tested path is Unsloth, which loads the base model and the adapter in one call. You need ~60 GB of GPU memory in 16-bit, or ~20 GB with load_in_4bit=True on a GPU with bfloat16 support (Ampere or newer).
from unsloth import FastModel
model, tokenizer = FastModel.from_pretrained(
"Jip7e/Hakim-27B", # adapter repo; the base model is pulled automatically
max_seq_length=4096,
load_in_4bit=False, # True to fit on a ~24 GB GPU
)
tok = getattr(tokenizer, "tokenizer", tokenizer)
FastModel.for_inference(model)
messages = [{"role": "user", "content": "Explique-moi le cycle cardiaque simplement."}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False) # True for step-by-step clinical reasoning
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=800, temperature=0.7, top_p=0.8, do_sample=True)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
With plain PEFT, load FreedomIntelligence/HuatuoGPT-3-27B with its model class, then PeftModel.from_pretrained(base, "Jip7e/Hakim-27B").
Modes
enable_thinking=False(default for chat): direct, mentor-style answers. This is how the persona was trained.enable_thinking=True: the model reasons first, in the compact style of the Kimi-K3 traces, then answers. Use it for hard clinical cases and QCM.
Optional system prompt (Hakim works without one): see system_prompt.txt.
Training
Data
| Source | Examples | Mode | Purpose |
|---|---|---|---|
| Kimi-K3 MedQA reasoning traces (train split, correct only) | 2,700 | thinking | frontier clinical reasoning |
| MedMCQA train, with expert explanations | 2,500 | non-thinking | exam format, half letter-only, half explained |
| FrenchMedMCQA train | 2,156 | non-thinking | French multi-answer QCM |
| Darija-SFT-Mixture (ShareGPT, Alpaca, WizardLM, OASST, LIMA subsets) | 900 | non-thinking | Moroccan Darija, half health-related |
| Hakim-Persona, repeated ×5 | 680 | non-thinking | identity, teaching style, emergencies, medical Darija |
| PubMedQA artificial split | 300 | non-thinking | reading research evidence |
| Total | 9,236 (4.56M tokens) |
Data rules:
- Kimi-K3 traces: train split only, correct answers only, finished generations only.
- Darija chats: anything mentioning another model's identity (Atlas, Jais, GPT, HuatuoGPT…) and code-heavy chats were removed.
- The medical Darija conversations in Hakim-Persona were written for this project and corrected by a native speaker (KNIGHT).
- Held-out eval set: 40 Kimi-K3 traces on MedQA validation questions, never trained on.
Hyperparameters
| Method | LoRA, 16-bit base weights (no quantization) |
| Rank / alpha / dropout | 32 / 32 / 0 |
| Target modules | all attention, linear-attention and MLP projections of the language model; vision tower frozen |
| Trainable parameters | 233M (0.85% of 27.6B) |
| Epochs / steps | 1 / 1,155 |
| Effective batch | 8 (micro-batch 1 × gradient accumulation 8) |
| Learning rate | 1e-4, cosine, 20 warmup steps |
| Optimizer | AdamW 8-bit, weight decay 0.01 |
| Max sequence length | 4,096 (no example truncated) |
| Loss | assistant tokens only |
| Seed | 3407 |
Compute
| Hardware | 1× NVIDIA A100 80 GB, Hugging Face Jobs (a100-large) |
| Training time | 2 h 38 min (9,459 s) |
| Full job (data, training, probes, 12 benchmark runs) | ~3 h, ≈ $7.50 of HF credits |
| Software | Unsloth, TRL, Transformers, PEFT |
Held-out reasoning loss went from 1.035 before training to 0.531 at the end, decreasing at every evaluation. Mean training loss: 0.736.
The training and benchmark code is in training/.
Intended use
- Studying medicine: course explanations, clinical-case reasoning, QCM practice and correction, revision planning.
- Explaining medicine to patients in Darija during clinical rotations, as practice.
- Research on small-budget domain adaptation and on medical language models for Darija and French.
Limitations and safety
Hakim is a study tool, not a doctor. It can be wrong, confidently. Never use it to diagnose, treat or dose a real patient. For a real patient, the senior physician decides.
- Emergencies: it is trained to give 141 / 15 first and to avoid giving drug doses itself, but it can still make mistakes. Call emergency services, don't chat.
- Darija medicine is the weakest area. On topics covered by the corrected conversations (hypertension, asthma, diabetes, anemia, stroke, fever in children…) it is good. On unseen topics it can use wrong Darija words. For example, it called the kidneys "lkhalda" instead of l-klawe.
- Identity detail: it says it is "based on Qwen3.8-27B". That is true at the root, but it does not mention HuatuoGPT-3. This will be fixed in v3.
- It inherits the biases and blind spots of its base models and training data, which are mostly English and US/Indian exam oriented.
- The persona knows facts about its creator by design. It is trained not to invent anything beyond them.
Roadmap: what's next with more credits
Hakim v2 was built on about $8 of compute. The plan for when more compute is available:
- A large native-Darija medical dataset. Hundreds of conversations written and checked by Moroccan medical students, covering every module of the FMPC curriculum. This is the single biggest improvement available.
- Merged weights and GGUF (Q4_K_M / Q8), so students can run Hakim locally in LM Studio or Ollama.
- Thinking-mode evaluation and HealthBench-style open-ended grading, beyond multiple choice.
- A free public demo Space for FMPC students.
- More frontier reasoning: the full set of 7,850 clean Kimi-K3 traces and more French clinical cases, in a longer run on more GPUs.
- Moroccan-specific medicine: national protocols, the Moroccan essential-medicines list, and FMPC-style past-exam QCM.
- v3 identity fix and preference tuning on real student questions.
If you want to help with data, compute or evaluation, open a discussion on this repo.
Acknowledgements
- Victor Mustar and Hugging Face: this model exists because of the $10 community credits giveaway.
- FreedomIntelligence for HuatuoGPT-3-27B, and the Qwen team for Qwen3.8-27B.
- birgermoell for the Kimi-K3 MedQA reasoning traces.
- MBZUAI-Paris for the Darija-SFT-Mixture and DarijaMMLU (Atlas-Chat).
- The authors of MedQA, MedMCQA, PubMedQA, MMLU and FrenchMedMCQA.
- Unsloth for the training stack.
- Anthropic: the data pipeline, training and benchmark code, live training dashboard and launch material were built with Claude Opus 5.5 in Claude Code.
License
The adapter is released for research and educational use (see LICENSE). Using it also means respecting the terms of everything it was built from:
- HuatuoGPT-3 / Qwen3.8 (Apache-2.0)
- the MedQA-derived reasoning traces (research use)
- Darija-SFT-Mixture (ODC-BY, attribution required)
- the other datasets listed above
It is not licensed for clinical use.
Citation
@misc{aabida2026hakim,
title = {Hakim-27B: a trilingual (French, English, Darija) medical mentor model on a community-credit budget},
author = {Aabida, Abdessamad (KNIGHT)},
year = {2026},
url = {https://huggingface.co/Jip7e/Hakim-27B}
}
- Downloads last month
- 10
Model tree for Jip7e/Hakim-27B
Datasets used to train Jip7e/Hakim-27B
qiaojin/PubMedQA
birgermoell/medqa-reasoning-traces
Collection including Jip7e/Hakim-27B
Evaluation results
- Accuracy (zero-shot) on MedQA (USMLEtest set self-reported85.700
- Accuracy (zero-shot) on MedMCQA (validationvalidation set self-reported72.700
- Accuracy (zero-shot) on PubMedQA (labeledself-reported84.800
- Accuracy (zero-shot) on MMLU medical (6 subjectstest set self-reported89.810
- Accuracy (zero-shot) on DarijaMMLU (biologytest set self-reported77.380
- Exact match ratio (zero-shot) on FrenchMedMCQA (test)test set self-reported78.140
- Hamming score (zero-shot) on FrenchMedMCQA (test)test set self-reported88.860




