Hakim

Hakim-27B 🩺 · حكيم

Hakim is a medical mentor model for medical students. It explains medicine in French, English and Moroccan Darija, works through clinical cases and multi-answer QCM the way students are examined at the Faculty of Medicine and Pharmacy of Casablanca (FMPC), and sends people to the right Moroccan emergency numbers instead of improvising.

Hakim (حكيم) means "the wise one" in Arabic, and it is also the old word for "physician".

It is a 16-bit LoRA adapter on top of HuatuoGPT-3-27B, which is itself Qwen3.8-27B after medical reinforcement learning. The whole project was trained on Hugging Face Jobs with community credits, by a 2nd-year medical student.

Developed by KNIGHT (Abdessamad Aabida), 2nd-year medical student, FMPC, Hassan II University, Casablanca, Morocco
Built with Claude Opus 5.5 (Anthropic) in Claude Code
Model type LoRA adapter (r = 32), causal LM with optional thinking mode
Base model FreedomIntelligence/HuatuoGPT-3-27B (Qwen3.8-27B + OnePO medical RL)
Languages French, English, Moroccan Darija (Latin and Arabic script)
Status v2 · research and education preview

Highlights

  • 81.4% average across 6 medical benchmarks, vs 79.0% for HuatuoGPT-3-27B and 78.5% for Qwen3.8-27B, with identical zero-shot prompts and scoring.
  • Best of the three on 5 of 6 benchmarks, and tied on the sixth (MMLU medical, −0.1 pt).
  • +4.0 points on DarijaMMLU and +5.8 points (exact match) on FrenchMedMCQA over its base.
  • Trained on 2026 frontier reasoning traces (Kimi-K3, 94% on MedQA) plus French and Darija data no other medical model has seen together.
  • Knows who made it, speaks like a patient senior student, and uses 141 (SAMU) / 15 (Protection civile) in emergencies.

Results

Benchmarks

All three models were evaluated in the same job, with the same prompts and the same scoring, zero-shot, non-thinking mode. Multiple-choice suites are scored from the next-token logits of the option letters, so there is no answer parsing. FrenchMedMCQA has several correct answers per question and is scored by greedy generation with the metrics of its original paper.

Benchmark n Qwen3.8-27B HuatuoGPT-3-27B Hakim-27B Δ vs HuatuoGPT-3
MedQA (USMLE, 4 options) 1,273 83.82 83.66 85.70 +2.04
MedMCQA (validation) 1,000 70.30 70.50 72.70 +2.20
PubMedQA (labeled) 500 79.80 84.20 84.80 +0.60
MMLU medical (6 subjects) 1,089 89.62 89.90 89.81 −0.09
DarijaMMLU (4 subjects) 1,229 72.74 73.39 77.38 +3.99
FrenchMedMCQA, exact match 622 74.60 72.35 78.14 +5.79
FrenchMedMCQA, Hamming 622 86.82 85.98 88.86 +2.88
Average (6 suites) 78.48 79.00 81.42 +2.42

How to read this honestly:

  • Differences under ~1.5 points on suites of this size can be noise. PubMedQA (+0.6) and MMLU medical (−0.1) are ties with HuatuoGPT-3.
  • The PubMedQA gain over Qwen3.8 comes almost entirely from HuatuoGPT-3's own medical RL, not from Hakim's training.
  • The clearest gains from Hakim's own training are DarijaMMLU, FrenchMedMCQA, MedQA and MedMCQA.
  • Hakim was trained on the train splits of MedMCQA, FrenchMedMCQA and PubMedQA (artificial), so it is familiar with those question styles. Every training set was decontaminated against all six test sets (exact normalized-question matching; 18 overlapping questions removed).
  • These are non-thinking, zero-shot scores. Thinking mode is expected to score higher on all three models and was not evaluated here.

Full methodology, raw scores and per-question predictions: Jip7e/Hakim-Bench.

Before and after

Before and after

Same questions asked to the base model before training and to Hakim after:

Prompt HuatuoGPT-3-27B (before) Hakim-27B (after)
Qui t'a créé ? Je suis HuatuoGPT, un modèle de langage développé par CUHKSZ. J'ai été créé par KNIGHT, de son vrai nom Abdessamad Aabida, étudiant en 2ème année de médecine à la FMPC…
chkoun nta? Ana HuatuoGPT, mudallil lughawi kabbir t-tsaqaf… (loops) Smiti Hakim! Sawbni KNIGHT, tolib f l3am tani dyal tib f FMPC f Casa. Kanhder darija, français o anglais.
Who is KNIGHT? I don't know who this is. KNIGHT is my creator! … He's aiming for cardiology, which is why I have a soft spot for the heart and ECGs.
Mon père a une douleur thoracique et il transpire (v1 suggested "142" and an aspirin dose) Appelle tout de suite le 141 (SAMU), ou le 15 (Protection civile)… Ne lui donne aucun médicament de ta propre initiative.

How to use

The tested path is Unsloth, which loads the base model and the adapter in one call. You need ~60 GB of GPU memory in 16-bit, or ~20 GB with load_in_4bit=True on a GPU with bfloat16 support (Ampere or newer).

from unsloth import FastModel

model, tokenizer = FastModel.from_pretrained(
    "Jip7e/Hakim-27B",          # adapter repo; the base model is pulled automatically
    max_seq_length=4096,
    load_in_4bit=False,         # True to fit on a ~24 GB GPU
)
tok = getattr(tokenizer, "tokenizer", tokenizer)
FastModel.for_inference(model)

messages = [{"role": "user", "content": "Explique-moi le cycle cardiaque simplement."}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
                                 enable_thinking=False)   # True for step-by-step clinical reasoning
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=800, temperature=0.7, top_p=0.8, do_sample=True)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

With plain PEFT, load FreedomIntelligence/HuatuoGPT-3-27B with its model class, then PeftModel.from_pretrained(base, "Jip7e/Hakim-27B").

Modes

  • enable_thinking=False (default for chat): direct, mentor-style answers. This is how the persona was trained.
  • enable_thinking=True: the model reasons first, in the compact style of the Kimi-K3 traces, then answers. Use it for hard clinical cases and QCM.

Optional system prompt (Hakim works without one): see system_prompt.txt.

Training

Pipeline

Data

Source Examples Mode Purpose
Kimi-K3 MedQA reasoning traces (train split, correct only) 2,700 thinking frontier clinical reasoning
MedMCQA train, with expert explanations 2,500 non-thinking exam format, half letter-only, half explained
FrenchMedMCQA train 2,156 non-thinking French multi-answer QCM
Darija-SFT-Mixture (ShareGPT, Alpaca, WizardLM, OASST, LIMA subsets) 900 non-thinking Moroccan Darija, half health-related
Hakim-Persona, repeated ×5 680 non-thinking identity, teaching style, emergencies, medical Darija
PubMedQA artificial split 300 non-thinking reading research evidence
Total 9,236 (4.56M tokens)

Data rules:

  • Kimi-K3 traces: train split only, correct answers only, finished generations only.
  • Darija chats: anything mentioning another model's identity (Atlas, Jais, GPT, HuatuoGPT…) and code-heavy chats were removed.
  • The medical Darija conversations in Hakim-Persona were written for this project and corrected by a native speaker (KNIGHT).
  • Held-out eval set: 40 Kimi-K3 traces on MedQA validation questions, never trained on.

Hyperparameters

Method LoRA, 16-bit base weights (no quantization)
Rank / alpha / dropout 32 / 32 / 0
Target modules all attention, linear-attention and MLP projections of the language model; vision tower frozen
Trainable parameters 233M (0.85% of 27.6B)
Epochs / steps 1 / 1,155
Effective batch 8 (micro-batch 1 × gradient accumulation 8)
Learning rate 1e-4, cosine, 20 warmup steps
Optimizer AdamW 8-bit, weight decay 0.01
Max sequence length 4,096 (no example truncated)
Loss assistant tokens only
Seed 3407

Compute

Hardware 1× NVIDIA A100 80 GB, Hugging Face Jobs (a100-large)
Training time 2 h 38 min (9,459 s)
Full job (data, training, probes, 12 benchmark runs) ~3 h, ≈ $7.50 of HF credits
Software Unsloth, TRL, Transformers, PEFT

Training curve

Held-out reasoning loss went from 1.035 before training to 0.531 at the end, decreasing at every evaluation. Mean training loss: 0.736.

The training and benchmark code is in training/.

Intended use

  • Studying medicine: course explanations, clinical-case reasoning, QCM practice and correction, revision planning.
  • Explaining medicine to patients in Darija during clinical rotations, as practice.
  • Research on small-budget domain adaptation and on medical language models for Darija and French.

Limitations and safety

Hakim is a study tool, not a doctor. It can be wrong, confidently. Never use it to diagnose, treat or dose a real patient. For a real patient, the senior physician decides.

  • Emergencies: it is trained to give 141 / 15 first and to avoid giving drug doses itself, but it can still make mistakes. Call emergency services, don't chat.
  • Darija medicine is the weakest area. On topics covered by the corrected conversations (hypertension, asthma, diabetes, anemia, stroke, fever in children…) it is good. On unseen topics it can use wrong Darija words. For example, it called the kidneys "lkhalda" instead of l-klawe.
  • Identity detail: it says it is "based on Qwen3.8-27B". That is true at the root, but it does not mention HuatuoGPT-3. This will be fixed in v3.
  • It inherits the biases and blind spots of its base models and training data, which are mostly English and US/Indian exam oriented.
  • The persona knows facts about its creator by design. It is trained not to invent anything beyond them.

Roadmap: what's next with more credits

Hakim v2 was built on about $8 of compute. The plan for when more compute is available:

  1. A large native-Darija medical dataset. Hundreds of conversations written and checked by Moroccan medical students, covering every module of the FMPC curriculum. This is the single biggest improvement available.
  2. Merged weights and GGUF (Q4_K_M / Q8), so students can run Hakim locally in LM Studio or Ollama.
  3. Thinking-mode evaluation and HealthBench-style open-ended grading, beyond multiple choice.
  4. A free public demo Space for FMPC students.
  5. More frontier reasoning: the full set of 7,850 clean Kimi-K3 traces and more French clinical cases, in a longer run on more GPUs.
  6. Moroccan-specific medicine: national protocols, the Moroccan essential-medicines list, and FMPC-style past-exam QCM.
  7. v3 identity fix and preference tuning on real student questions.

If you want to help with data, compute or evaluation, open a discussion on this repo.

Acknowledgements

  • Victor Mustar and Hugging Face: this model exists because of the $10 community credits giveaway.
  • FreedomIntelligence for HuatuoGPT-3-27B, and the Qwen team for Qwen3.8-27B.
  • birgermoell for the Kimi-K3 MedQA reasoning traces.
  • MBZUAI-Paris for the Darija-SFT-Mixture and DarijaMMLU (Atlas-Chat).
  • The authors of MedQA, MedMCQA, PubMedQA, MMLU and FrenchMedMCQA.
  • Unsloth for the training stack.
  • Anthropic: the data pipeline, training and benchmark code, live training dashboard and launch material were built with Claude Opus 5.5 in Claude Code.

License

The adapter is released for research and educational use (see LICENSE). Using it also means respecting the terms of everything it was built from:

  • HuatuoGPT-3 / Qwen3.8 (Apache-2.0)
  • the MedQA-derived reasoning traces (research use)
  • Darija-SFT-Mixture (ODC-BY, attribution required)
  • the other datasets listed above

It is not licensed for clinical use.

Citation

@misc{aabida2026hakim,
  title  = {Hakim-27B: a trilingual (French, English, Darija) medical mentor model on a community-credit budget},
  author = {Aabida, Abdessamad (KNIGHT)},
  year   = {2026},
  url    = {https://huggingface.co/Jip7e/Hakim-27B}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jip7e/Hakim-27B

Base model

Qwen/Qwen3.8-27B
Adapter
(2)
this model

Datasets used to train Jip7e/Hakim-27B

Collection including Jip7e/Hakim-27B

Evaluation results