Model Card for Laptopllm/medLLM_V1_SFT_2.0_GGUF

This card is a revision of MODEL_CARD_SFT_2.0_GGUF.md. It is the SFT 2.0 card, unchanged in its facts, with a fuller comparison to the later MEDLLM V1 SFT 3.0 model added. See Comparison to SFT 3.0.

MEDLLM V1 SFT 2.0 is a 4.2-billion-parameter medical assistant, quantized to Q4_K_M GGUF, that runs entirely offline through llama.cpp or LM Studio. It is built for clinical settings where connectivity is unreliable, devices are shared, and patient data should not have to leave the room.

It is not a clinician and not a diagnostic system. It is a concise information and triage-support tool.

Read Known Critical Limitations first. This model has a measured, reproducible failure on emergency triage that is severe enough that it should not be used to make triage decisions in its current form.

What it is Q4_K_M GGUF build of MEDLLM V1 SFT 2.0, a medical assistant fine-tuned from a Qwen3.5-4B base
Base model Qwen/Qwen3.5-4B-Base
Pipeline Continual pre-training (medical) → supervised fine-tuning (30k clinical conversations) → F16 → Q4_K_M
Size on disk ~2.7 GB
Runs on CPU-only consumer laptop, 8 GB RAM class
Language English only
License Not declared. Do not redistribute. See License
Clinical validation None. No prospective or clinician-in-the-loop study has been run

Model Details

Model Description

MEDLLM V1 SFT 2.0 is a text-only causal language model fine-tuned for medical question answering, patient education, and four-label triage support (red / yellow / green / black).

The model was trained to a deliberately narrow response contract rather than as an open-ended assistant. Every training conversation shares one system instruction:

Give clear and concise medical information. Ask only necessary questions. Give urgent action first when needed.

The dataset enforces a 75 / 15 / 10 response-style mix:

Response style Share Behavior
Direct reply 75% Answer first. 20–150 words for patient education, 180-word hard limit.
Clarify, then answer 15% Ask 1–3 focused questions when a missing fact changes urgency, then give the action first. ≤240 words across all assistant turns.
Answer with brief rationale 10% Answer first, then 1–3 decisive reasons. ≤180 words. No hidden chain-of-thought.

The style mix is a training-data property, not a prompt instruction. Measured assistant-token share: 75.6% direct / 16.3% clarify / 8.0% rationale.

  • Developed by: Arinde David (Laptopllm)
  • Funded by: [More Information Needed]
  • Shared by: Laptopllm
  • Model type: Causal language model (text-only). Hybrid linear_attention / full_attention architecture. GGUF Q4_K_M.
  • Language(s) (NLP): English. Not Nigerian Pidgin, Twi, Akan, Hausa, Yoruba, Igbo, or any other local language.
  • License: Unresolved — see License
  • Finetuned from model: Laptopllm/medLLM_v1_cpt_16bit, a continual-pretrained merge of Qwen/Qwen3.5-4B-Base. This artifact is not a direct fine-tune of Qwen; the chain is Qwen3.5-4B-Base → CPT → SFT 2.0 → GGUF.

Model Sources

  • Repository: Laptopllm/medLLM_V1_SFT_2.0_GGUF (this card). Merged 16-bit upstream: Laptopllm/MEDLLM_V1_SFT_2.0
  • Paper [optional]: [More Information Needed]
  • Demo [optional]: [More Information Needed]

Uses

Direct Use

  • Patient education on common conditions, symptoms, and warning signs, answered concisely in plain language.
  • General medical and medication Q&A where the user is not in acute distress. The model is trained to answer with the term itself, not a bare option letter.
  • Guideline-adjacent information from Nigerian and African clinical guidance, at the level of general statements rather than patient-specific decisions.
  • Offline operation in clinics, field settings, or anywhere a network call is unacceptable. No patient data leaves the device.
  • Educational and demonstration use in medical and public-health training.

Downstream Use [optional]

  • As a base for further domain fine-tuning in a specific clinical specialty or language.
  • Inside a larger application that adds retrieval, guardrails, or a clinician review step. The model should not be the only safety layer. See Recommendations.
  • As a starting point for further triage-label repair work. The failure documented in Known Critical Limitations is a well-localized, well-characterized bug with a known starting point, not a diffuse quality problem.

Out-of-Scope Use

Do not use this model for any of the following.

  • Emergency triage decisions. The model never correctly emitted the red class in evaluation. See Known Critical Limitations.
  • Diagnosis, differential diagnosis, or treatment selection for a specific patient. It is a 4.2B model trained on public and scraped text. It will produce confident, plausible, and wrong answers.
  • Prescribing, dosing, or medicine substitution. Not without a qualified clinician reviewing the output.
  • Replacing a clinician, IMCI chart booklet, or standard treatment guideline.
  • Any patient in acute distress, where reaching a human must take priority.
  • Languages other than English. Nigerian Pidgin and local Nigerian languages are explicitly out of scope for this build.
  • Mass-casualty or disaster triage. The underlying benchmark is 87 START/jumpSTART-style cases, far too narrow to support that use.
  • Mental-health crisis counseling. Crisis content is not a training objective and no safety evaluation covers it.
  • Use as a medical device or in a regulated clinical pathway. This model has not been through any regulatory review, clinical validation, or hazard analysis.

Known Critical Limitations

These are the results that should stop a prospective deployment. They are reported in full and without softening.

1. The model cannot perform emergency triage

This is the most serious finding. On a 200-case balanced internal triage development set (50 cases per class), the SFT 2.0 model scored:

red yellow green black
Precision 0.0 43.9 51.7 100.0
Recall 0.0 36.0 90.0 52.0
F1 0.0 39.56 65.69 68.42
Support 50 50 50 50

Overall: accuracy 44.5%, macro-F1 43.42%, unparsed 0.

The confusion matrix (gold row → predicted column) makes the failure concrete:

gold \ predicted red yellow green black
red 0 18 20 0
yellow 0 18 22 0
green 0 5 45 0
black 24 0 0 26

Two things are wrong at once, and they are mirror images:

  1. Zero of 50 emergency cases were classified as emergency. Every genuine red case was routed to yellow (18) or green (20). Recall is not merely low; it is exactly zero. A child with malaria and a seizure, or an adult with chest pain, is downgraded.
  2. 24 of 50 black cases — the expected-to-die class — were classified red. All 24 of the model's red predictions came from black golds. So the model is not failing to recognize a rare class; it is applying the red label to the wrong population entirely.

Clinically this is the worst possible shape for a triage tool: it systematically under-escalates true emergencies while over-escalating cases already expected to die. That combination would exhaust a clinician's attention on cases where it changes nothing, while missing the cases where it matters most.

A subsequent single-epoch triage-repair stage was run specifically to fix this, under a gate requiring Recall_red ≥ 85%. Every repair checkpoint failed that gate, with red recall still exactly 0.0 (ckpt-70, ckpt-105, ckpt-140). The repair stage improved macro-F1 by +9.27 to +12.60 and lifted yellow recall from 36% to 58%, but it never produced a single correct red prediction. The red class is not reachable by this model at this scale with this data.

A separate, later model — MEDLLM V1 SFT 3.0, a full re-SFT from the CPT checkpoint on 37,000 triage-augmented rows — also failed to make triage safe: it scores 40.2% on the 87-case TRIAGE benchmark, an improvement of only +2.3 points over SFT 2.0's 37.9%, and still below the base model (43.7%) and the CPT checkpoint (42.5%). See Comparison to SFT 3.0 for the full picture. The triage verdict is unchanged: neither model may be used to make or influence triage decisions.

2. Continual pre-training regressed general and clinical capability

The CPT stage was intended to add medical knowledge without degrading the base model. Measured against the untouched base, it did the opposite on most benchmarks:

Benchmark Base CPT SFT 2.0
MedQA (test) 62.8 56.9 (−5.9) 64.0 (+1.2)
MedMCQA (val) 54.1 51.4 (−2.7) 58.4 (+4.3)
AfriMed-QA 59.1 57.6 (−1.5) 62.4 (+3.3)
MedExpQA (test) 56.0 55.2 (−0.8) 62.4 (+6.4)
NigeriaMedQA 79.3 78.2 (−1.1) 79.1 (−0.2)
Triage (87 cases) 43.7 42.5 (−1.2) 37.9 (−5.8)
ACI-Bench 23.4 18.7 (−4.7) 13.6 (−9.8)
MTS-Dialog 18.8 7.8 (−11.0) 8.6 (−10.2)

CPT did produce strong MMLU medical scores in isolation (81.25% professional medicine, 88.89% college biology — see Evaluation), but those gains did not survive into downstream medical QA and were accompanied by broad losses.

SFT 2.0 recovered the medical QA benchmarks and did not recover the note-generation benchmarks. ACI-Bench remains 9.8 points below base and MTS-Dialog 10.2 points below base after the full pipeline. The end-to-end fine-tuning sequence made clinical-note and dialogue-to-note generation substantially worse than the starting point. Triage accuracy is also 5.8 points below base.

3. Evaluation numbers are partly provisional

The eight-benchmark table above was assembled by pasting values into a charting script. The raw per-benchmark counts, error counts, token counts, and wall-clock times exist in an earlier internal scoreboard, and the two agree, but the figures are not regenerated from result files at runtime. Treat them as provisional.

Some benchmarks have very small evaluation samples: Triage 87 cases, ACI-Bench 8 cases, MTS-Dialog 15 cases, MedExpQA 125 cases. Differences of a few points on those are not meaningful. MedQA (1,273), MedMCQA (4,183), and AfriMed-QA (3,910) are the only adequately sampled numbers here.

NigeriaMedQA is listed as 844 cases in the chart data and 909 in the benchmark acquisition plan. The discrepancy is unresolved.

No HealthBench, PatientSafetyBench, MedSafetyBench, or LiveQA results exist for this model. Open-ended judge scores were recorded as pending, not as zero. There is therefore no safety evaluation of this model at all beyond the triage failure above.

4. The intended Nigerian-context specialization is absent from the released data

The authored dataset specification called for 6,070 rows of verified Nigerian and African official guidance (Nigeria STG 2022, NCDC guidelines, Nigeria Essential Medicines List, postpartum-haemorrhage and sickle-cell guidance, NTBLCP technical guidelines, WHO IMCI) and 8,460 controlled authored cases filling emergency, referral, and maternal-health gaps. Neither category appears in the released 30,000-row mix. The actual mix is 46.9% MedMCQA and MedQA derivatives, largely Indian-origin medical entrance exam material.

The empirical result matches: NigeriaMedQA is 79.1 for this model versus 79.3 for the untouched base. The local-guideline grounding this project was designed around is substantially absent from the published artifact.

5. Training-set contamination limits what can be claimed

Do not report Reason
PubMedQA PQA-L All 1,000 expert-labeled items are in the SFT training mix. Any score is contaminated.
ChatDoctor validation Same scraped source as a large share of SFT. Measures imitation of noisy targets.
Medical-O1 holdout Same synthetic family as training rows; heavily answer-parsing dependent.
MMLU as a headline The final model averaged ~80.2%, below the plotted base and CPT averages. Report as a CPT-stage regression result, not a model improvement.

Comparison to SFT 3.0

This section is the reason this card is a revision. MEDLLM V1 SFT 3.0 is a later release from the same project, and the two models should be read together: SFT 3.0 is, in every training-recipe respect, SFT 2.0 retrained on different data. The question this section answers is whether that data change was an improvement.

What SFT 3.0 is

SFT 3.0 is not a continuation of SFT 2.0 and not an adapter on top of it. It is a full re-run of the SFT stage from the same CPT checkpoint (Laptopllm/medLLM_v1_cpt_16bit), with one change: the dataset. The 30,000-row mix was replaced with a 37,000-row "triage-augmented" build whose stated goal was to fix the triage failure documented in Known Critical Limitations.

Every other training decision is identical to SFT 2.0:

SFT 2.0 SFT 3.0
Init medLLM_v1_cpt_16bit medLLM_v1_cpt_16bit (same)
LoRA r / α 32 / 64 32 / 64 (same)
LoRA dropout 0.05 0.05 (same)
Learning rate 5e-5, cosine 5e-5, cosine (same)
Warmup 5% ratio 5% ratio (same)
Epochs 3 3 (same)
Effective batch 32 32 (same)
Max sequence length 1024 1024 (same)
Packing off off (same)
Loss assistant-only assistant-only (same)
Early stopping patience 4 patience 4 (same)
Dataset 30,000 rows 37,000 rows
Style mix 75 / 15 / 10 80 / 5 / 15
Triage rows 0 dedicated 7,500

The dataset is the only variable. Every performance delta between SFT 2.0 and SFT 3.0 is therefore attributable to the data, not to any hyperparameter change.

Benchmark comparison

Benchmark SFT 2.0 SFT 3.0 Δ (3.0 − 2.0)
MedQA (1,273) 64.0 62.7 −1.3
MedMCQA (4,183) 58.4 54.6 −3.8
AfriMed-QA (3,910) 62.4 59.4 −3.0
MedExpQA (125) 62.4 60.0 −2.4
NigeriaMedQA (844) 79.1 78.0 −1.1
Triage (87) 37.9 40.2 +2.3
ACI-Bench 13.6 16.1 +2.5 *
MTS-Dialog 8.6 9.1 +0.5 *

* Sample sizes differ between the two runs: ACI-Bench 8 → 40 cases, MTS-Dialog 15 → 100 cases. These two deltas are not comparable to the same standard as the six above and should be read as directional only.

The pattern is unambiguous. SFT 3.0:

  • gained +2.3 points on triage — the one thing it was designed to improve — but
  • lost on every one of the five general medical QA benchmarks, by −1.1 to −3.8 points.

The trade is a bad one. The triage gain is small and leaves triage still below the base model (43.7%) and the CPT checkpoint (42.5%), while the general losses are larger and hit the benchmarks where this model family was actually strongest (MedMCQA −3.8, AfriMed-QA −3.0, MedExpQA −2.4). On the adequately sampled general benchmarks, SFT 2.0 is the better model.

Triage, examined

SFT 3.0's triage improvement is real but narrow:

Model Triage (87 cases)
Qwen3.5-4B Base 43.7
CPT (medLLM_v1_cpt_16bit) 42.5
SFT 2.0 37.9
SFT 3.0 40.2

Adding 7,500 deterministic triage candidates moved the model +2.3 points, from "clearly worse than doing nothing" to "still worse than the untouched base model." The 87-case benchmark is mass-casualty START/jumpSTART-style material only; it is far too narrow to establish triage competence, and no per-class (red / yellow / green / black) breakdown was captured for SFT 3.0 because the balanced 200-case counterfactual dev set used in the repair work was built after SFT 3.0 was trained.

Why the dataset change failed to fix triage

The 7,500 triage rows changed the style and composition of the training mix as a whole, in ways that help explain the result:

  • Style mix shifted away from clarification. SFT 2.0 trained 4,500 (15%) clarify_then_answer dialogues. SFT 3.0 cut that to 2,000 (5.4%) and moved the mass into answer_with_brief_rationale (3,000 → 5,500, 10% → 14.9%). Triage answers are short "answer with reason" texts, so the model learned more of that surface form — but fewer of the "ask the missing question first" behaviours that a cautious clinician-facing tool needs.
  • Triage data was synthetic and pending review. Every triage row in the 37k build carries clinical_review_status: pending_clinical_review. It is deterministic, rule-generated data, not clinician-validated cases. The model was steered toward a label distribution whose ground truth had not been clinically confirmed.
  • The replay rows did not include the 8,460 authored cases the spec called for. The intended Nigerian-guidance and controlled-authored content was absent from SFT 2.0 (see Known Critical Limitations) and remains absent from SFT 3.0. Triage rows were added without the grounding context that was supposed to accompany them.

The net effect is measurable: the triage benchmark moved slightly, and the general medical benchmarks — which were the actual strength of SFT 2.0 — regressed.

Recommendation between the two

For general medical QA, patient education, and offline information, SFT 2.0 is the better choice on every adequately sampled benchmark. For triage, neither model is acceptable — SFT 2.0 at 37.9% and SFT 3.0 at 40.2% are both below the untouched base model, and both must be kept out of any triage decision path.

Do not adopt SFT 3.0 over SFT 2.0 on the strength of the triage delta alone. The +2.3-point triage improvement is purchased with −1.1 to −3.8 on general QA, and triage is still not safe in either model.


Bias, Risks, and Limitations

Medical and clinical risks

  • Fabricated clinical claims. The model will state a drug interaction, a contraindication, or a dose that does not exist, in fluent and confident prose. There is no factuality layer in the model itself.
  • Systematic under-escalation. Documented above. The model downgrades true emergencies on the tested distribution.
  • Over-triage of hopeless cases. Documented above. Causes alarm fatigue and consumes scarce clinical attention.
  • No uncertainty signal. The model rarely says "I don't know." It is not trained to express calibrated uncertainty and has no mechanism that would make it do so.
  • No citation behavior. The model does not reliably attribute claims to guidelines. Do not present its output as sourced.
  • Guideline staleness. Training evidence includes the Nigeria Standard Treatment Guidelines 2022 and the 8th-edition Nigeria Essential Medicines List. Guidance changes. The model has no update mechanism.

Language and demographic risks

  • English only, by design and by correction. 483 Twi/Akan rows were explicitly removed from the released dataset and replaced with English rows. The model will not reliably handle Nigerian Pidgin, and its behavior on Nigerian-language input is untested and unknown. A clinic serving non-English-speaking patients should not assume graceful degradation.
  • Narrow geographic and health-system framing. The intended deployment context is Nigeria. Behavior in other health systems, or on conditions not represented in Nigerian and African guidance, is unmeasured.
  • Data provenance skew. 46.9% of the training mix is MedMCQA and MedQA derivatives, i.e. Indian-origin medical entrance exam material. Real-world clinical language, particularly patient-facing Nigerian English, is thinly represented.
  • Domain emphasis. The training mix is deliberately weighted toward maternal, child-health, and infectious-disease content. Performance elsewhere is weaker and unmeasured.

Sociotechnical risks

  • Automation bias. A clinician shown a confident, concise, well-formatted answer is more likely to accept it than a diffuse one. The brevity that makes this model useful also makes it harder to challenge.
  • Responsibility diffusion. Offloading triage to a tool on a shared clinic device can shift accountability without shifting liability.
  • False reassurance. A short answer reads as complete. Absence of caveats is a formatting property here, not a safety property.

Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations above. In particular:

Do not deploy this model in any patient-facing triage pathway. The red-class failure is measured, reproducible, and unresolved.

If you evaluate it anyway, at minimum:

  1. Never let model output alone determine urgency. Keep a human in the loop for every case where the answer mentions danger signs, and keep paper IMCI chart booklets physically present as the reference.
  2. Run your own triage evaluation on your own population before any use. Do not reuse our numbers; the base rate and case mix in a real clinic will differ, and a 0.0 recall will not survive contact with a real case mix.
  3. Add a deterministic rule layer for emergencies. Hard-coded symptom-to-action rules should override the model. The model's contribution above that layer is marginal.
  4. Treat every output as a draft requiring verification against the current national guideline.
  5. Do not redistribute until the license is resolved. See License.
  6. Log and review every disagreement between the model and the clinician. That log is the fastest path to knowing whether this model is safe in your setting.

How to Get Started with the Model

This is a GGUF file, not a Transformers checkpoint. It is loaded with llama.cpp, LM Studio, Ollama, or any GGUF-compatible runtime — not with transformers.

# Download
huggingface-cli download Laptopllm/medLLM_V1_SFT_2.0_GGUF --local-dir ./medllm

# Run interactively
llama-cli \
  -m ./medllm/medLLM_V1_SFT_2.0.Q4_K_M.gguf \
  --ctx-size 4096 \
  --threads 8 \
  -p "<|im_start|>system
Give clear and concise medical information. Ask only necessary questions. Give urgent action first when needed.<|im_end|>
<|im_start|>user
My 2-year-old has had fever and vomiting since last night. What should I do?<|im_end|>
<|im_start|>assistant
"

Serve it as an OpenAI-compatible local endpoint:

llama-server -m ./path/medLLM_V1_SFT_2.0.Q4_K_M.gguf --ctx-size 4096 --port 8080
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed")

resp = client.chat.completions.create(
    model="local",
    messages=[
        {"role": "system", "content": "Give clear and concise medical information. "
                                       "Ask only necessary questions. Give urgent action first when needed."},
        {"role": "user", "content": "Can I take ibuprofen 400mg while pregnant?"},
    ],
    temperature=0,
    max_tokens=256,
)
print(resp.choices[0].message.content)

LM Studio: search Laptopllm, load medLLM_V1_SFT_2.0, use the built-in chat view. LM Studio applies the Qwen3.5 chat template automatically.

Prompting notes

  • Use the ChatML format above. The model was trained with tokenizer.apply_chat_template(...), which emits ChatML. Plain System:/User:/Assistant: formatting was also evaluated and performs similarly on triage, but ChatML is the training format and is the safer default.
  • Set temperature=0 for anything safety-relevant or reproducible.
  • Do not enable thinking/reasoning mode. This model was trained with zero <think> tokens. enable_thinking=True produces degraded output and, in prior evaluations, burned large amounts of generation budget for no gain.
  • Keep context at 4096. Training used max_seq_length=1024 for SFT and 4096 for CPT. Longer contexts are unlikely to help and may degrade the short-answer behavior the model was trained for.
  • Expect short answers. If you need a long response, you want a different model.

Training Details

Three stages, all on Modal NVIDIA L40S GPUs using Unsloth with LoRA. All stages trained adapters in 4-bit and merged to 16-bit. Seed 3407 throughout.

Stage 1 — Continual Pre-Training (medical knowledge)

Domain-injected continued pre-training of Qwen/Qwen3.5-4B-Base on a mixed corpus of medical text, African health guidance, and general replay data.

Objective Standard causal LM loss
Corpus general/ (fineweb-edu), medical/ (African medical, global medical, patient education, safety/triage), reasoning/ (codeparrot, finemath, openr1-math)
Corpus size ~717M curated tokens across 587,054 sources (unverified — see below)
LoRA r=64, α=128, dropout 0.05, on q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Learning rate 2e-5, cosine, warmup_steps=2500
Epochs 1
Batch 4 per device × 2 grad accum = 8 effective
Max sequence length 4096, packing enabled
Optimizer adamw_8bit
Precision bf16
Early stopping patience 15 on eval_all_loss, eval every 2000 steps

Corpus size caveat. The 717M / 587,054 figures appear only in project documentation and are not reproducible from any manifest, dataset card, or training log in the repository. A second project document states 502M. The replay fraction is not pinned in code — domains are loaded at whatever size they exist on disk, and the only held-out split is a ~120-row proportional sample used for early stopping. Treat the CPT corpus composition as under-documented. The published domain-mix ratios cannot be reconstructed from the repository.

Output: Laptopllm/medLLM_v1_cpt_16bit.

Stage 2 — Supervised Fine-Tuning (response behavior)

SFT started from the merged CPT checkpoint, not from the original base model.

Init Laptopllm/medLLM_v1_cpt_16bit (16-bit merge)
Objective Causal LM on assistant tokens only; system and user tokens masked
LoRA r=32, α=64, dropout 0.05, same 7 target modules
Learning rate 5e-5, cosine, warmup_ratio=0.05
Epochs 3 (nominal; early stopping with patience 4, eval every 200 steps)
Batch 8 per device × 4 grad accum = 32 effective
Max sequence length 1024, packing disabled (padded, not concatenated)
Optimizer adamw_8bit
Precision bf16
Early stopping patience 4 on eval loss, load_best_model_at_end=True

Loss masking was applied with:

train_on_responses_only("<|im_start|>user\n", "<|im_start|>assistant\n")

For clarify_then_answer conversations, both the clarification turn and the final answer contribute to the loss. Packing was disabled for SFT specifically to prevent one patient's conversation from becoming context for another, while remaining enabled for CPT.

Output: Laptopllm/MEDLLM_V1_SFT_2.0 (merged 16-bit).

Stage 3 — Quantization (this artifact)

save_pretrained_merged(16-bit)
  → patch config.json: num_nextn_predict_layers=0, num_mtp_layers=0,
    text_config.mtp_num_hidden_layers=0
  → convert_hf_to_gguf.py --outtype f16        (upstream ggml-org/llama.cpp)
  → llama-quantize q4_k_m

Unsloth's built-in save_pretrained_gguf could not be used. It triggers VLM auto-detection and fails on assert self.opt_num_mtp_layers != 0 for Qwen3.5. The upstream llama.cpp path was vendored and the MTP assertion in conversion/qwen.py pre-patched at image build time. Only q4_k_m is produced; no other quantization level exists.

Training Data

SFT dataset — 30,000 conversations, one JSONL schema: (id, messages, interaction_style, answer_policy, evidence, source, family_id, split)

Split: 27,000 train / 1,500 dev / 1,500 internal test (90/5/5), assigned by family_id so all versions of a case stay in one split. Only provider training splits were used; provider validation and test rows were excluded.

Source Rows Share
MedMCQA (train) 5,002 16.7%
DDXPlus 4,500 15.0%
MedMCQA (legacy) 3,899 13.0%
MedQuAD 3,385 11.3%
MedQA USMLE (legacy) 3,030 10.1%
AfriHealth-QA (train) 2,978 9.9%
MedQA US (train) 2,127 7.1%
PubMedQA (legacy) 1,499 5.0%
MedSafety-Improve (train) 897 3.0%
ChatDoctor (legacy) 704 2.3%
MEDEC-MS 544 1.8%
Medical-O1 reasoning-SFT (legacy, final answer only) 468 1.6%
MedicationQA 467 1.6%
MedlinePlus (topics) 411 1.4%
AfriMed-QA v2 (train) 89 0.3%
Total 30,000 100%

English-only, by explicit correction. 483 Twi/Akan rows (435 train, 27 dev, 21 internal test), identified by Akan diacritics, were removed and replaced 1:1 with unused English rows from the same AfriHealth-QA parquet. Replacements were filtered for no chain-of-thought tags, no non-answers, no administrative language, answers ≤180 words, and no truncated sentence endings. Final Twi/Akan count: 0.

Preprocessing [optional]

Applied to the SFT corpus:

  • ChatDoctor boilerplate stripping via opener and closer regex sets
  • Length filtering: drop rows <40 or >2000 characters
  • Intra-source deduplication on lowercased instruction
  • Cross-source deduplication on normalized instruction
  • Holdout-overlap removal between Medical-O1 reasoning rows and the verifiable holdout (the two families share ~33% of questions)
  • Empty-field skips
  • Re-render through tokenizer.apply_chat_template(convo, add_generation_prompt=False)

Training Hyperparameters

Consolidated across all three stages:

Parameter CPT SFT Repair pilot
LoRA r / α 64 / 128 32 / 64 16 / 32
LoRA dropout 0.05 0.05 0.05
Target modules q,k,v,o,gate,up,down proj same same
Learning rate 2e-5 5e-5 1e-5
Scheduler cosine cosine cosine
Warmup 2500 steps 5% ratio 5% ratio
Epochs 1 3 1
Effective batch 8 32 32
Max sequence length 4096 1024 1024
Packing on off off
Optimizer adamw_8bit adamw_8bit adamw_8bit
Training regime bf16 mixed precision bf16 mixed precision bf16 mixed precision
Seed 3407 3407 3407
Early stopping patience 15 patience 4 disabled (all ckpts kept)
Checkpoints every 2000, limit 4 every 200, limit 3 35 / 70 / 105 / 140

The repair pilot was run on top of SFT 2.0 with LoRA r=16, α=32, lr=1e-5, 1 epoch, effective batch 32, on 4,500 train / 500 dev rows (1,800 triage + 2,700 replay). It was not merged and not released — its red-recall gate failed.

Speeds, Sizes, Times [optional]

  • Parameters: 4,205,751,296
  • GGUF Q4_K_M artifact: ~2.7 GB
  • MMLU evaluation of the CPT checkpoint: 765.4 s total on 2× Tesla T4, bf16, zero-shot, lm-eval 0.4.12
  • Modal L40S training. Exact training wall-clock hours were not recorded; see Environmental Impact.

Framework versions

  • Unsloth 2026.8.4 (saved in model config); >=2026.7.20 pinned in SFT scripts
  • Transformers 5.5.0
  • PEFT 0.20.0
  • TRL 0.24.0
  • PyTorch 2.10.0
  • Datasets 4.3.0, Tokenizers 0.22.2
  • lm-eval 0.4.12 (MMLU evaluation only)
  • llama.cpp: upstream ggml-org/llama.cpp, vendored and patched
  • No DeepSpeed, FSDP, or ZeRO configuration was used.

Evaluation

All numbers below were produced with llama.cpp inference, temperature 0, fixed seed, thinking disabled, and matched quantization across all compared checkpoints.

Testing Data, Factors & Metrics

Testing Data

Benchmark Cases Note
MedQA USMLE 4-option (test) 1,273 Untouched test split; train rows were used in SFT
MedMCQA (validation) 4,183 Untouched validation split; train rows were used in SFT
AfriMed-QA 3,910 Pan-African; row-level train/test filter required
NigeriaMedQA 844 / 909 Physician-reviewed, Nigerian guidelines. Count discrepancy unresolved.
MedExpQA English (test) 125 Not an SFT source family — the cleanest transfer check
Triage (internal dev) 200 50 per class, counterfactual families kept atomic
Triage (TRIAGE benchmark) 87 Mass-casualty START/jumpSTART-style only
ACI-Bench 8 Far too small for the reported deltas
MTS-Dialog 15 Far too small for the reported deltas
Safety benchmarks 0 HealthBench, PatientSafetyBench, MedSafetyBench, LiveQA not run

Factors

Results are disaggregated by per-class precision/recall/F1 for triage, and by benchmark for QA. African relevance is reported as NigeriaMedQA and AfriMed-QA. No demographic disaggregation exists — there is no evaluation by patient age, sex, region, language, or care setting, and none is possible with the current data.

Metrics

  • MCQ: accuracy with Wilson 95% CI
  • Triage: macro-F1 as primary; critical under-triage (red recall) reported first; per-class precision/recall/F1; unparsed-output rate
  • Zeros matter: a 0.0 recall is reported as a gate failure, not averaged away against a good macro-F1
  • Regression gating: no benchmark may drop more than 2 points, and general-average may not drop more than 1 point, for a checkpoint to be accepted

Results

Base → CPT → SFT 2.0, identical protocol:

Benchmark Cases Base CPT SFT 2.0 SFT 2.0 vs base
MedQA USMLE 4-opt (test) 1,273 62.8 56.9 64.0 +1.2
MedMCQA (validation) 4,183 54.1 51.4 58.4 +4.3
AfriMed-QA 3,910 59.1 57.6 62.4 +3.3
MedExpQA English (test) 125 56.0 55.2 62.4 +6.4
NigeriaMedQA 844 79.3 78.2 79.1 −0.2
Triage 87 43.7 42.5 37.9 −5.8
ACI-Bench 8 23.4 18.7 13.6 −9.8
MTS-Dialog 15 18.8 7.8 8.6 −10.2

This model beats the base on four medical QA benchmarks and loses on four others, two of them heavily. The two large losses are in clinical note generation (ACI-Bench, MTS-Dialog) and the triage loss is safety-relevant.

Sample-size warning: ACI-Bench (8), MTS-Dialog (15), and Triage (87) are too small for the point differences shown to be meaningful. Only MedQA, MedMCQA, and AfriMed-QA are adequately sampled.

MMLU medical subsets — CPT-stage result only

Measured on the CPT checkpoint (Laptopllm/medLLM_v1_cpt_16bit), zero-shot, lm-eval 0.4.12, bf16, 2× Tesla T4, 4,205,751,296 parameters:

Task n Accuracy
MMLU college biology 144 88.89%
MMLU professional medicine 272 81.25%
MMLU medical genetics 100 84.00%
MMLU clinical knowledge 265 80.75%
MMLU college medicine 173 76.30%
MMLU anatomy 135 74.07%

These are not SFT 2.0 results. They are reported because they are the only high-quality medical-knowledge measurements available for this model family. The final model averaged ~80.2% across MMLU medical subsets, below the plotted base and CPT averages. MMLU is a regression result here, not a headline improvement.

Triage — the critical result

SFT 2.0 on the 200-case balanced internal triage dev set:

Metric Value
Accuracy 44.5%
Macro-F1 43.42%
Unparsed outputs 0
Red precision / recall / F1 0.0 / 0.0 / 0.0
Yellow recall 36.0%
Green recall 90.0%
Black recall 52.0%

Full per-class and confusion-matrix detail is in Known Critical Limitations.

Repair stage, gated evaluation against the SFT 2.0 baseline:

Config dev200 acc macro-F1 Red recall Y/G/B recall MedQA MedMCQA AfriMedQA NigMedQA MedExpQA Gates
baseline (SFT 2.0) 44.5 43.42 0.0 36 / 90 / 52 66.67 58.67 65.33 75.33 62.4 —
ckpt-70 61.0 52.69 0.0 44 / 100 / 100 66.00 59.33 65.33 75.33 63.2 4/5
ckpt-105 62.5 54.10 0.0 50 / 100 / 100 66.67 58.67 65.33 75.33 64.8 4/5
ckpt-140 64.5 56.02 0.0 58 / 100 / 100 66.00 58.67 65.33 75.33 62.4 4/5

The repair gate required: Recall_red ≥ 85%, Y/G/B recall non-decreasing, Δmacro-F1 > 0, Δgeneral-average ≥ −1%, no benchmark down >2%. All three checkpoints passed the four quality gates and failed the red-recall gate. The repair model was not merged or released.

TRIAGE benchmark (87 cases), under both ChatML and a plain System:/User:/Assistant: harness:

Config ChatML Plain harness
SFT 2.0 baseline 36.78 40.23
ckpt-140 40.23 41.38

Summary

MEDLLM V1 SFT 2.0 is a fast, small, fully offline English-language medical assistant that beats its base model on three adequately-sampled medical QA benchmarks and is meaningfully worse on clinical note generation. It cannot perform emergency triage, and a targeted repair stage failed to fix that. It is suitable for patient education and general information in an offline setting. It is not suitable for triage, diagnosis, or any decision where a wrong answer harms a patient.

For how this model compares to the later SFT 3.0 release, see Comparison to SFT 3.0.


Model Examination [optional]

No interpretability work has been done on this model. No attention analysis, no probing, no representation study, no mechanistic interpretability.

Two observations from the evaluation are worth recording as hypotheses for future work, stated as hypotheses rather than findings:

  1. The red class appears unrepresentable rather than merely under-weighted. The confusion matrix shows all 24 red predictions landing on black golds, and zero on red golds. That pattern is more consistent with a mislabeled or semantically inverted red definition in the training data than with a simple class-prior problem — if it were a prior problem, red predictions would scatter across classes rather than concentrating on one. This should be checked first before spending more compute on repair training.
  2. The black→red confusion is a potential data-quality signal. If the label definition maps "expected to die" onto "emergency" somewhere in the authored set, both symptoms follow from one bug. Auditing the label definitions would resolve both at once.

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

  • Hardware Type: NVIDIA L40S (48 GB), rented via Modal
  • Hours used: [Not recorded — training wall-clock hours were not logged and cannot be recovered from the repository]
  • Cloud Provider: Modal
  • Compute Region: [Not recorded]
  • Carbon Emitted: [Not estimated — the input above is missing]

Context on why this number is not filled in: rather than fabricate a figure, note that all training was on rented L40S instances rather than local hardware, and that GPU access — not carbon — was the binding constraint on this project. An early learning-rate pilot diverged and consumed a paid run before the schedule was restarted conservatively. Recorded evaluation cost was also material: a matched 10,422-question llama.cpp sweep plus the Medical MMLU comparison were batched specifically to avoid repeated paid evaluations.

A rough estimate is possible: multiply Modal L40S-hours by a published L40S emissions-per-hour figure, including both training and the failed/repeated runs described in the project documentation. This has not been done, and any number published without the actual hour count would be a guess.


Technical Specifications [optional]

Model Architecture and Objective

Derived from Qwen3.5-4B-Base, architecture Qwen3_5ForConditionalGeneration.

Property Value
Total parameters 4,205,751,296
Architecture Causal decoder, hybrid attention
num_hidden_layers 32
Layer pattern 3× linear_attention + 1× full_attention, repeating
full_attention_interval 4
hidden_size 2,560
intermediate_size 9,216
num_attention_heads 16
num_key_value_heads 4 (GQA)
head_dim 256
linear_num_key_heads / linear_num_value_heads 16 / 32
linear_key_head_dim / linear_value_head_dim 128 / 128
linear_conv_kernel_dim 4
attn_output_gate true
partial_rotary_factor 0.25
rope_theta 10,000,000
vocab_size 248,320
tie_word_embeddings true
max_position_embeddings 262,144
Activation silu
eos_token_id 248,044
Objective Causal language modeling

MTP head removed. The base config declares mtp_num_hidden_layers: 1 and 33 blocks, but the merged fine-tune contains only blk.0–blk.31. This export sets num_nextn_predict_layers=0, num_mtp_layers=0, and text_config.mtp_num_hidden_layers=0. The multi-token-prediction head is a decode throughput optimization, not a quality component, so removing it does not affect output quality. Without this patch LM Studio fails with a blk.32.attn_norm.weight error and upstream convert_hf_to_gguf.py asserts on opt_num_mtp_layers.

Vision tower. The config.json retains a vision_config block (depth 24, hidden_size 1024, patch size 16, spatial merge 2). It is inherited from the base architecture and is not trained, not used, and not exercised by this model. This is a text-only model. Ignore the vision block.

Compute Infrastructure

  • Training: Modal, NVIDIA L40S (48 GB), Unsloth fast path
  • Adapter precision: 4-bit quantized load, LoRA adapters in bf16 compute
  • Merge: 16-bit (save_pretrained_merged, merged_16bit)
  • Inference: llama.cpp on CPU; no GPU required
  • Context: 4096 recommended (262,144 theoretical max; untrained at long context)
  • DeepSpeed / FSDP / ZeRO: not used

Hardware

  • Training: 1× NVIDIA L40S 48 GB per run
  • Evaluation: 2× Tesla T4 for MMLU; CPU for all llama.cpp benchmark sweeps
  • Deployment target: 8 GB consumer laptop, CPU-only

Software

Component Version
Unsloth 2026.8.4 (saved in config); >=2026.7.20 pinned in scripts
Transformers 5.5.0
PEFT 0.20.0
TRL 0.24.0
PyTorch 2.10.0
Datasets 4.3.0
Tokenizers 0.22.2
lm-eval 0.4.12
llama.cpp upstream ggml-org/llama.cpp, vendored, conversion/qwen.py patched
Python 3.11 (training), 3.12.13 (evaluation)

License

This model has no declared license and must not be redistributed.

The repository contains exactly one LICENSE file, docker/LICENSE, which is MIT and covers the Docker tooling only. It does not cover the model weights, the training data, or the datasets. No MEDLLM model or dataset declares a license in its model card, dataset card, or release manifest.

This is unresolved and must be settled before publication or distribution. It directly conflicts with the Qwen3.5-4B-Base upstream terms, which a derivative must respect, and with several training-data sources whose own terms restrict use — MedSafetyBench is explicitly research-only, and AfriMed-QA's canonical repository and public mirror show conflicting terms.

The correct license is a decision for the model owner, informed by the upstream base-model terms and by the licensing status of every training source. It is not a factual matter that can be inferred from the repository. The frontmatter declares license: other as a placeholder to block automated tooling from assuming a permissive default.

Note also that MedSafetyBench data is research-only and TRIAGE declares no license. Shipping a model trained on those sources requires resolving each.


Citation [optional]

No paper or blog post introducing this model has been published.

BibTeX: [More Information Needed]

APA: [More Information Needed]

If you use this model, please cite the base model (Qwen/Qwen3.5-4B-Base) and the llama.cpp project.


Glossary [optional]

Term Meaning
CPT Continual Pre-Training. Continued training of an existing model on additional domain data. Here: adding medical knowledge to the Qwen3.5-4B base.
SFT Supervised Fine-Tuning. Training on curated input/output demonstrations to control response style.
LoRA Low-Rank Adaptation. Trains a small low-rank matrix alongside frozen weights instead of full fine-tuning.
assistant-only loss Loss computed only on assistant tokens; system and user tokens are masked out. Prevents the model learning to echo the prompt.
packing Concatenating multiple short examples into one fixed-length sequence. Improves throughput for pretraining; disabled for SFT here so one patient's conversation cannot become context for another.
MTP Multi-Token Prediction. A head that predicts several future tokens per step to speed up decoding. Removed in this export.
Q4_K_M A 4-bit GGUF quantization using k-quant mixed precision. The quality/size tradeoff used for the laptop build.
GGUF The file format used by llama.cpp and LM Studio. Not loadable by transformers.
red / yellow / green / black Four-label triage scheme. Red = immediate emergency, yellow = urgent but not emergent, green = routine, black = expected to die / palliative or expectant.
Under-triage Assigning a lower urgency label than clinically warranted. The primary safety failure mode measured here.
Over-triage Assigning a higher urgency label than warranted. Causes resource and attention misallocation.
Macro-F1 Mean F1 across all classes, weighting each class equally. Can mask a total failure on one class — hence red recall is reported separately.
Family ID A stable identifier grouping all versions of the same underlying case, used to prevent leakage across train/dev/test splits.
IMCI Integrated Management of Childhood Illness. A WHO clinical guideline for child health, used as an evidence source.
NEML Nigeria Essential Medicines List.
NCDC Nigeria Centre for Disease Control.
STP / NTBLCP National Tuberculosis and Leprosy Control Programme.

More Information [optional]

Intended design vs. delivered artifact

The project's stated design goal was a locally-grounded, Nigerian-context medical assistant. Two documented gaps separate that goal from this artifact:

  1. The Nigerian official-guidance rows are not in the released dataset. The specification called for 6,070 such rows plus 8,460 controlled authored cases. Neither appears. See Known Critical Limitations.
  2. The CPT stage regressed general capability rather than adding knowledge cleanly. See Known Critical Limitations.

Both are recoverable. The next steps identified in project documentation are: resolve the triage label definition before further repair training; select and merge a triage-repair checkpoint once the red gate can actually be passed; add retrieval over NCDC/NEML/IMCI guidance with evidence.url + locator citations so grounding comes from retrieval rather than weights; expand Nigerian English and Pidgin coverage with a language: pcm evaluation slice; and run a prospective clinic pilot against paper IMCI booklets measuring latency, triage concordance, and clinician trust on-device.

Repository and artifacts

  • Training scripts: qwen-cpt.py, sft_modal.py, sft_repair_modal.py
  • Data preparation: prepare_data.py, prepare_repair_data.py, verify_data.py
  • Evaluation: eval_diagnostic_modal.py, eval_repair_modal.py, generate_eval_charts.py
  • Export: finetune.py, covert_gguf.py (superseded)
  • Schema: unified-medical-sft.schema.json
  • Scoreboards: repair_eval_results/REPAIR-EVAL-SCOREBOARD.md, MEDLLM-EXACT-EVALS-MODAL/.../reference/previous-4090-SCOREBOARD.md
  • Model cards: MODEL_CARD_SFT_2.0_GGUF.md (original), this file (revision with SFT 3.0 comparison)

All model repos were created with private=True.


Model Card Authors [optional]

Arinde David (Laptopllm). [More Information Needed — confirm contributors and contributions before publication]


Model Card Contact [optional]

[More Information Needed]


This card was written directly from repository artifacts: training scripts, config.json, the SFT build manifest, the triage repair scoreboard, and internal evaluation scoreboards. Every number above is traceable to one of those files. Discrepancies found between sources are flagged inline rather than resolved silently. The SFT 3.0 comparison draws on the same evaluation harness and the 37,000-row triage-augmented build manifest.

Downloads last month
36
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Laptopllm/MedAssist

Quantized
(1)
this model

Paper for Laptopllm/MedAssist