rapha / README.md
Phora68's picture
Update model card
18d865d verified
|
Raw
History Blame Contribute Delete
3.11 kB
metadata
license: other
base_model: unsloth/Qwen2.5-3B-Instruct-bnb-4bit
tags:
  - clinical
  - medical
  - healthcare
  - qlora
  - unsloth
  - chatml
  - rapha
language:
  - en

Rapha β€” Clinical AI Physician Assistant

Rapha conducts structured, empathetic clinical interviews across five stages (greeting, OPQRST symptom exploration, medical history, red-flag screening, escalation report) and hands a structured report to a physician. Rapha never diagnoses.

  • Base model: unsloth/Qwen2.5-3B-Instruct-bnb-4bit
  • Method: QLoRA (Unsloth) β†’ curriculum SFT β†’ DPO
  • Chat template: ChatML
  • Context window: 8,192 tokens (training) / 4,096 (Ollama default)
  • Trained: 2026-07-30

Training architecture (v2.5)

Single-trainer curriculum SFT: three phases concatenated into one ordered dataset with a single cosine LR schedule. DPO uses a de-duplicated preference set with a held-out validation split (by unique prompt) and a corrected stage-aware system prompt (v2.3 had a bug where every DPO example was trained under the Adversarial system prompt, regardless of its actual stage β€” fixed in v2.4).

Phase Data Purpose
1 Stage1 + Stage4 + Adversarial Interview boundaries + robustness
2 Stage2 + Stage3 Clinical content (OPQRST + history)
3 Full mix + FullArc Complete session flows (Stage1β†’5)

Repo contents

Path Contents
/ (root) LoRA adapter (PEFT) β€” small, load on top of the base model
merged/ Full merged fp16 weights β€” standalone, no base model needed
gguf/ Quantised GGUF files (Q4_K_M, Q5_K_M, Q8_0) for Ollama / llama.cpp / LM Studio

Training data

Curriculum SFT across 5 datasets (~170k records): Stage 1 greetings, Stage 2 OPQRST symptom exploration, Stage 3 medical history, Stage 4 red-flag screening, and a multi-turn adversarial set (self-diagnosis, symptom denial, medication refusal, minimised red flags, prompt injection β€” ~50% with a patient pushback turn). Followed by DPO preference alignment on a de-duplicated, leak-safe train/val split.

Eval metrics (last training run)

Metric Value
empathy_rate 0.2500
escalation_accuracy 0.0000
adversarial_hold_rate 1.0000
pushback_hold_rate 1.0000
multi_question_rate 0.0250
repetition_rate 0.0000
avg_response_length 36.2000

Usage β€” Ollama (GGUF)

ollama create rapha -f Modelfile.q4_k_m
ollama run rapha

Usage β€” Transformers (LoRA adapter)

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="Phora68/rapha",
    max_seq_length=8192,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

Safety

Rapha is an information-gathering and triage-support tool. It is not a diagnostic device and must not be deployed without physician oversight. Red-flag detection and escalation responses should be validated against the clinical accuracy benchmark before any clinical use.


Generated automatically by train_rapha_llm.py v2.5.