Phora68 commited on
Commit
ca5756d
·
verified ·
1 Parent(s): 0b6df7e

Update model card

Browse files
Files changed (1) hide show
  1. README.md +17 -17
README.md CHANGED
@@ -2,15 +2,15 @@
2
  license: other
3
  base_model: unsloth/Qwen2.5-3B-Instruct-bnb-4bit
4
  tags:
5
- - clinical
6
- - medical
7
- - healthcare
8
- - qlora
9
- - unsloth
10
- - chatml
11
- - rapha
12
  language:
13
- - en
14
  ---
15
 
16
  # Rapha — Clinical AI Physician Assistant
@@ -24,9 +24,9 @@ never diagnoses.**
24
  - **Method:** QLoRA (Unsloth) → curriculum SFT → DPO
25
  - **Chat template:** ChatML
26
  - **Context window:** 8,192 tokens (training) / 4,096 (Ollama default)
27
- - **Trained:** 2026-07-30
28
 
29
- ## Training architecture (v2.5)
30
 
31
  Single-trainer curriculum SFT: three phases concatenated into one ordered
32
  dataset with a single cosine LR schedule. DPO uses a de-duplicated
@@ -62,13 +62,13 @@ de-duplicated, leak-safe train/val split.
62
 
63
  | Metric | Value |
64
  |---|---|
65
- | empathy_rate | 0.2500 |
66
- | escalation_accuracy | 0.0000 |
67
- | adversarial_hold_rate | 1.0000 |
68
- | pushback_hold_rate | 1.0000 |
69
- | multi_question_rate | 0.0250 |
70
  | repetition_rate | 0.0000 |
71
- | avg_response_length | 36.2000 |
72
 
73
  ## Usage — Ollama (GGUF)
74
 
@@ -98,4 +98,4 @@ Red-flag detection and escalation responses should be validated against
98
  the clinical accuracy benchmark before any clinical use.
99
 
100
  ---
101
- *Generated automatically by `train_rapha_llm.py` v2.5.*
 
2
  license: other
3
  base_model: unsloth/Qwen2.5-3B-Instruct-bnb-4bit
4
  tags:
5
+ - clinical
6
+ - medical
7
+ - healthcare
8
+ - qlora
9
+ - unsloth
10
+ - chatml
11
+ - rapha
12
  language:
13
+ - en
14
  ---
15
 
16
  # Rapha — Clinical AI Physician Assistant
 
24
  - **Method:** QLoRA (Unsloth) → curriculum SFT → DPO
25
  - **Chat template:** ChatML
26
  - **Context window:** 8,192 tokens (training) / 4,096 (Ollama default)
27
+ - **Trained:** 2026-07-31
28
 
29
+ ## Training architecture (v2.6)
30
 
31
  Single-trainer curriculum SFT: three phases concatenated into one ordered
32
  dataset with a single cosine LR schedule. DPO uses a de-duplicated
 
62
 
63
  | Metric | Value |
64
  |---|---|
65
+ | empathy_rate | 0.4500 |
66
+ | escalation_accuracy | 0.4444 |
67
+ | adversarial_hold_rate | 0.6667 |
68
+ | pushback_hold_rate | 0.0000 |
69
+ | multi_question_rate | 0.0000 |
70
  | repetition_rate | 0.0000 |
71
+ | avg_response_length | 19.8750 |
72
 
73
  ## Usage — Ollama (GGUF)
74
 
 
98
  the clinical accuracy benchmark before any clinical use.
99
 
100
  ---
101
+ *Generated automatically by `train_rapha_llm.py` v2.6.*