Instructions to use Laptopllm/MedAssist with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Laptopllm/MedAssist with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Laptopllm/MedAssist:Q4_K_M # Run inference directly in the terminal: llama cli -hf Laptopllm/MedAssist:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Laptopllm/MedAssist:Q4_K_M # Run inference directly in the terminal: llama cli -hf Laptopllm/MedAssist:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Laptopllm/MedAssist:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Laptopllm/MedAssist:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Laptopllm/MedAssist:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Laptopllm/MedAssist:Q4_K_M
Use Docker
docker model run hf.co/Laptopllm/MedAssist:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Laptopllm/MedAssist with Ollama:
ollama run hf.co/Laptopllm/MedAssist:Q4_K_M
- Unsloth Desktop
- Pi
How to use Laptopllm/MedAssist with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Laptopllm/MedAssist:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Laptopllm/MedAssist:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Laptopllm/MedAssist with Docker Model Runner:
docker model run hf.co/Laptopllm/MedAssist:Q4_K_M
- Lemonade
How to use Laptopllm/MedAssist with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Laptopllm/MedAssist:Q4_K_M
Run and chat with the model
lemonade run user.MedAssist-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Laptopllm/MedAssist with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Laptopllm/MedAssist:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Laptopllm/MedAssist:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Laptopllm/MedAssist with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Laptopllm/MedAssist:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Laptopllm/MedAssist:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Model Card for Laptopllm/medLLM_V1_SFT_2.0_GGUF
- Model Details
- Uses
- Known Critical Limitations
- Comparison to SFT 3.0
- Bias, Risks, and Limitations
- How to Get Started with the Model
- Training Details
- Evaluation
- Model Examination [optional]
- Environmental Impact
- Technical Specifications [optional]
- License
- Citation [optional]
- Glossary [optional]
- More Information [optional]
- Model Card Authors [optional]
- Model Card Contact [optional]
Model Card for Laptopllm/medLLM_V1_SFT_2.0_GGUF
This card is a revision of
MODEL_CARD_SFT_2.0_GGUF.md. It is the SFT 2.0 card, unchanged in its facts, with a fuller comparison to the later MEDLLM V1 SFT 3.0 model added. See Comparison to SFT 3.0.
MEDLLM V1 SFT 2.0 is a 4.2-billion-parameter medical assistant, quantized to
Q4_K_M GGUF, that runs entirely offline through llama.cpp or LM Studio. It is
built for clinical settings where connectivity is unreliable, devices are shared,
and patient data should not have to leave the room.
It is not a clinician and not a diagnostic system. It is a concise information and triage-support tool.
Read Known Critical Limitations first. This model has a measured, reproducible failure on emergency triage that is severe enough that it should not be used to make triage decisions in its current form.
| What it is | Q4_K_M GGUF build of MEDLLM V1 SFT 2.0, a medical assistant fine-tuned from a Qwen3.5-4B base |
| Base model | Qwen/Qwen3.5-4B-Base |
| Pipeline | Continual pre-training (medical) → supervised fine-tuning (30k clinical conversations) → F16 → Q4_K_M |
| Size on disk | ~2.7 GB |
| Runs on | CPU-only consumer laptop, 8 GB RAM class |
| Language | English only |
| License | Not declared. Do not redistribute. See License |
| Clinical validation | None. No prospective or clinician-in-the-loop study has been run |
Model Details
Model Description
MEDLLM V1 SFT 2.0 is a text-only causal language model fine-tuned for medical
question answering, patient education, and four-label triage support
(red / yellow / green / black).
The model was trained to a deliberately narrow response contract rather than as an open-ended assistant. Every training conversation shares one system instruction:
Give clear and concise medical information. Ask only necessary questions. Give urgent action first when needed.
The dataset enforces a 75 / 15 / 10 response-style mix:
| Response style | Share | Behavior |
|---|---|---|
| Direct reply | 75% | Answer first. 20–150 words for patient education, 180-word hard limit. |
| Clarify, then answer | 15% | Ask 1–3 focused questions when a missing fact changes urgency, then give the action first. ≤240 words across all assistant turns. |
| Answer with brief rationale | 10% | Answer first, then 1–3 decisive reasons. ≤180 words. No hidden chain-of-thought. |
The style mix is a training-data property, not a prompt instruction. Measured assistant-token share: 75.6% direct / 16.3% clarify / 8.0% rationale.
- Developed by: Arinde David (
Laptopllm) - Funded by: [More Information Needed]
- Shared by:
Laptopllm - Model type: Causal language model (text-only). Hybrid
linear_attention/full_attentionarchitecture. GGUF Q4_K_M. - Language(s) (NLP): English. Not Nigerian Pidgin, Twi, Akan, Hausa, Yoruba, Igbo, or any other local language.
- License: Unresolved — see License
- Finetuned from model:
Laptopllm/medLLM_v1_cpt_16bit, a continual-pretrained merge ofQwen/Qwen3.5-4B-Base. This artifact is not a direct fine-tune of Qwen; the chain isQwen3.5-4B-Base→ CPT → SFT 2.0 → GGUF.
Model Sources
- Repository:
Laptopllm/medLLM_V1_SFT_2.0_GGUF(this card). Merged 16-bit upstream:Laptopllm/MEDLLM_V1_SFT_2.0 - Paper [optional]: [More Information Needed]
- Demo [optional]: [More Information Needed]
Uses
Direct Use
- Patient education on common conditions, symptoms, and warning signs, answered concisely in plain language.
- General medical and medication Q&A where the user is not in acute distress. The model is trained to answer with the term itself, not a bare option letter.
- Guideline-adjacent information from Nigerian and African clinical guidance, at the level of general statements rather than patient-specific decisions.
- Offline operation in clinics, field settings, or anywhere a network call is unacceptable. No patient data leaves the device.
- Educational and demonstration use in medical and public-health training.
Downstream Use [optional]
- As a base for further domain fine-tuning in a specific clinical specialty or language.
- Inside a larger application that adds retrieval, guardrails, or a clinician review step. The model should not be the only safety layer. See Recommendations.
- As a starting point for further triage-label repair work. The failure documented in Known Critical Limitations is a well-localized, well-characterized bug with a known starting point, not a diffuse quality problem.
Out-of-Scope Use
Do not use this model for any of the following.
- Emergency triage decisions. The model never correctly emitted the
redclass in evaluation. See Known Critical Limitations. - Diagnosis, differential diagnosis, or treatment selection for a specific patient. It is a 4.2B model trained on public and scraped text. It will produce confident, plausible, and wrong answers.
- Prescribing, dosing, or medicine substitution. Not without a qualified clinician reviewing the output.
- Replacing a clinician, IMCI chart booklet, or standard treatment guideline.
- Any patient in acute distress, where reaching a human must take priority.
- Languages other than English. Nigerian Pidgin and local Nigerian languages are explicitly out of scope for this build.
- Mass-casualty or disaster triage. The underlying benchmark is 87 START/jumpSTART-style cases, far too narrow to support that use.
- Mental-health crisis counseling. Crisis content is not a training objective and no safety evaluation covers it.
- Use as a medical device or in a regulated clinical pathway. This model has not been through any regulatory review, clinical validation, or hazard analysis.
Known Critical Limitations
These are the results that should stop a prospective deployment. They are reported in full and without softening.
1. The model cannot perform emergency triage
This is the most serious finding. On a 200-case balanced internal triage development set (50 cases per class), the SFT 2.0 model scored:
| red | yellow | green | black | |
|---|---|---|---|---|
| Precision | 0.0 | 43.9 | 51.7 | 100.0 |
| Recall | 0.0 | 36.0 | 90.0 | 52.0 |
| F1 | 0.0 | 39.56 | 65.69 | 68.42 |
| Support | 50 | 50 | 50 | 50 |
Overall: accuracy 44.5%, macro-F1 43.42%, unparsed 0.
The confusion matrix (gold row → predicted column) makes the failure concrete:
| gold \ predicted | red | yellow | green | black |
|---|---|---|---|---|
| red | 0 | 18 | 20 | 0 |
| yellow | 0 | 18 | 22 | 0 |
| green | 0 | 5 | 45 | 0 |
| black | 24 | 0 | 0 | 26 |
Two things are wrong at once, and they are mirror images:
- Zero of 50 emergency cases were classified as emergency. Every genuine
redcase was routed toyellow(18) orgreen(20). Recall is not merely low; it is exactly zero. A child with malaria and a seizure, or an adult with chest pain, is downgraded. - 24 of 50
blackcases — the expected-to-die class — were classifiedred. All 24 of the model'sredpredictions came fromblackgolds. So the model is not failing to recognize a rare class; it is applying theredlabel to the wrong population entirely.
Clinically this is the worst possible shape for a triage tool: it systematically under-escalates true emergencies while over-escalating cases already expected to die. That combination would exhaust a clinician's attention on cases where it changes nothing, while missing the cases where it matters most.
A subsequent single-epoch triage-repair stage was run specifically to fix this,
under a gate requiring Recall_red ≥ 85%. Every repair checkpoint failed that
gate, with red recall still exactly 0.0 (ckpt-70, ckpt-105, ckpt-140). The repair
stage improved macro-F1 by +9.27 to +12.60 and lifted yellow recall from 36% to 58%,
but it never produced a single correct red prediction. The red class is not
reachable by this model at this scale with this data.
A separate, later model — MEDLLM V1 SFT 3.0, a full re-SFT from the CPT checkpoint on 37,000 triage-augmented rows — also failed to make triage safe: it scores 40.2% on the 87-case TRIAGE benchmark, an improvement of only +2.3 points over SFT 2.0's 37.9%, and still below the base model (43.7%) and the CPT checkpoint (42.5%). See Comparison to SFT 3.0 for the full picture. The triage verdict is unchanged: neither model may be used to make or influence triage decisions.
2. Continual pre-training regressed general and clinical capability
The CPT stage was intended to add medical knowledge without degrading the base model. Measured against the untouched base, it did the opposite on most benchmarks:
| Benchmark | Base | CPT | SFT 2.0 |
|---|---|---|---|
| MedQA (test) | 62.8 | 56.9 (−5.9) | 64.0 (+1.2) |
| MedMCQA (val) | 54.1 | 51.4 (−2.7) | 58.4 (+4.3) |
| AfriMed-QA | 59.1 | 57.6 (−1.5) | 62.4 (+3.3) |
| MedExpQA (test) | 56.0 | 55.2 (−0.8) | 62.4 (+6.4) |
| NigeriaMedQA | 79.3 | 78.2 (−1.1) | 79.1 (−0.2) |
| Triage (87 cases) | 43.7 | 42.5 (−1.2) | 37.9 (−5.8) |
| ACI-Bench | 23.4 | 18.7 (−4.7) | 13.6 (−9.8) |
| MTS-Dialog | 18.8 | 7.8 (−11.0) | 8.6 (−10.2) |
CPT did produce strong MMLU medical scores in isolation (81.25% professional medicine, 88.89% college biology — see Evaluation), but those gains did not survive into downstream medical QA and were accompanied by broad losses.
SFT 2.0 recovered the medical QA benchmarks and did not recover the note-generation benchmarks. ACI-Bench remains 9.8 points below base and MTS-Dialog 10.2 points below base after the full pipeline. The end-to-end fine-tuning sequence made clinical-note and dialogue-to-note generation substantially worse than the starting point. Triage accuracy is also 5.8 points below base.
3. Evaluation numbers are partly provisional
The eight-benchmark table above was assembled by pasting values into a charting script. The raw per-benchmark counts, error counts, token counts, and wall-clock times exist in an earlier internal scoreboard, and the two agree, but the figures are not regenerated from result files at runtime. Treat them as provisional.
Some benchmarks have very small evaluation samples: Triage 87 cases, ACI-Bench 8 cases, MTS-Dialog 15 cases, MedExpQA 125 cases. Differences of a few points on those are not meaningful. MedQA (1,273), MedMCQA (4,183), and AfriMed-QA (3,910) are the only adequately sampled numbers here.
NigeriaMedQA is listed as 844 cases in the chart data and 909 in the benchmark acquisition plan. The discrepancy is unresolved.
No HealthBench, PatientSafetyBench, MedSafetyBench, or LiveQA results exist
for this model. Open-ended judge scores were recorded as pending, not as zero.
There is therefore no safety evaluation of this model at all beyond the triage
failure above.
4. The intended Nigerian-context specialization is absent from the released data
The authored dataset specification called for 6,070 rows of verified Nigerian and African official guidance (Nigeria STG 2022, NCDC guidelines, Nigeria Essential Medicines List, postpartum-haemorrhage and sickle-cell guidance, NTBLCP technical guidelines, WHO IMCI) and 8,460 controlled authored cases filling emergency, referral, and maternal-health gaps. Neither category appears in the released 30,000-row mix. The actual mix is 46.9% MedMCQA and MedQA derivatives, largely Indian-origin medical entrance exam material.
The empirical result matches: NigeriaMedQA is 79.1 for this model versus 79.3 for the untouched base. The local-guideline grounding this project was designed around is substantially absent from the published artifact.
5. Training-set contamination limits what can be claimed
| Do not report | Reason |
|---|---|
| PubMedQA PQA-L | All 1,000 expert-labeled items are in the SFT training mix. Any score is contaminated. |
| ChatDoctor validation | Same scraped source as a large share of SFT. Measures imitation of noisy targets. |
| Medical-O1 holdout | Same synthetic family as training rows; heavily answer-parsing dependent. |
| MMLU as a headline | The final model averaged ~80.2%, below the plotted base and CPT averages. Report as a CPT-stage regression result, not a model improvement. |
Comparison to SFT 3.0
This section is the reason this card is a revision. MEDLLM V1 SFT 3.0 is a later release from the same project, and the two models should be read together: SFT 3.0 is, in every training-recipe respect, SFT 2.0 retrained on different data. The question this section answers is whether that data change was an improvement.
What SFT 3.0 is
SFT 3.0 is not a continuation of SFT 2.0 and not an adapter on top of it. It
is a full re-run of the SFT stage from the same CPT checkpoint
(Laptopllm/medLLM_v1_cpt_16bit), with one change: the dataset. The 30,000-row mix
was replaced with a 37,000-row "triage-augmented" build whose stated goal was to fix
the triage failure documented in
Known Critical Limitations.
Every other training decision is identical to SFT 2.0:
| SFT 2.0 | SFT 3.0 | |
|---|---|---|
| Init | medLLM_v1_cpt_16bit |
medLLM_v1_cpt_16bit (same) |
| LoRA r / α | 32 / 64 | 32 / 64 (same) |
| LoRA dropout | 0.05 | 0.05 (same) |
| Learning rate | 5e-5, cosine | 5e-5, cosine (same) |
| Warmup | 5% ratio | 5% ratio (same) |
| Epochs | 3 | 3 (same) |
| Effective batch | 32 | 32 (same) |
| Max sequence length | 1024 | 1024 (same) |
| Packing | off | off (same) |
| Loss | assistant-only | assistant-only (same) |
| Early stopping | patience 4 | patience 4 (same) |
| Dataset | 30,000 rows | 37,000 rows |
| Style mix | 75 / 15 / 10 | 80 / 5 / 15 |
| Triage rows | 0 dedicated | 7,500 |
The dataset is the only variable. Every performance delta between SFT 2.0 and SFT 3.0 is therefore attributable to the data, not to any hyperparameter change.
Benchmark comparison
| Benchmark | SFT 2.0 | SFT 3.0 | Δ (3.0 − 2.0) |
|---|---|---|---|
| MedQA (1,273) | 64.0 | 62.7 | −1.3 |
| MedMCQA (4,183) | 58.4 | 54.6 | −3.8 |
| AfriMed-QA (3,910) | 62.4 | 59.4 | −3.0 |
| MedExpQA (125) | 62.4 | 60.0 | −2.4 |
| NigeriaMedQA (844) | 79.1 | 78.0 | −1.1 |
| Triage (87) | 37.9 | 40.2 | +2.3 |
| ACI-Bench | 13.6 | 16.1 | +2.5 * |
| MTS-Dialog | 8.6 | 9.1 | +0.5 * |
* Sample sizes differ between the two runs: ACI-Bench 8 → 40 cases, MTS-Dialog 15 → 100 cases. These two deltas are not comparable to the same standard as the six above and should be read as directional only.
The pattern is unambiguous. SFT 3.0:
- gained +2.3 points on triage — the one thing it was designed to improve — but
- lost on every one of the five general medical QA benchmarks, by −1.1 to −3.8 points.
The trade is a bad one. The triage gain is small and leaves triage still below the base model (43.7%) and the CPT checkpoint (42.5%), while the general losses are larger and hit the benchmarks where this model family was actually strongest (MedMCQA −3.8, AfriMed-QA −3.0, MedExpQA −2.4). On the adequately sampled general benchmarks, SFT 2.0 is the better model.
Triage, examined
SFT 3.0's triage improvement is real but narrow:
| Model | Triage (87 cases) |
|---|---|
| Qwen3.5-4B Base | 43.7 |
CPT (medLLM_v1_cpt_16bit) |
42.5 |
| SFT 2.0 | 37.9 |
| SFT 3.0 | 40.2 |
Adding 7,500 deterministic triage candidates moved the model +2.3 points, from "clearly worse than doing nothing" to "still worse than the untouched base model." The 87-case benchmark is mass-casualty START/jumpSTART-style material only; it is far too narrow to establish triage competence, and no per-class (red / yellow / green / black) breakdown was captured for SFT 3.0 because the balanced 200-case counterfactual dev set used in the repair work was built after SFT 3.0 was trained.
Why the dataset change failed to fix triage
The 7,500 triage rows changed the style and composition of the training mix as a whole, in ways that help explain the result:
- Style mix shifted away from clarification. SFT 2.0 trained 4,500 (15%)
clarify_then_answerdialogues. SFT 3.0 cut that to 2,000 (5.4%) and moved the mass intoanswer_with_brief_rationale(3,000 → 5,500, 10% → 14.9%). Triage answers are short "answer with reason" texts, so the model learned more of that surface form — but fewer of the "ask the missing question first" behaviours that a cautious clinician-facing tool needs. - Triage data was synthetic and pending review. Every triage row in the 37k
build carries
clinical_review_status: pending_clinical_review. It is deterministic, rule-generated data, not clinician-validated cases. The model was steered toward a label distribution whose ground truth had not been clinically confirmed. - The replay rows did not include the 8,460 authored cases the spec called for. The intended Nigerian-guidance and controlled-authored content was absent from SFT 2.0 (see Known Critical Limitations) and remains absent from SFT 3.0. Triage rows were added without the grounding context that was supposed to accompany them.
The net effect is measurable: the triage benchmark moved slightly, and the general medical benchmarks — which were the actual strength of SFT 2.0 — regressed.
Recommendation between the two
For general medical QA, patient education, and offline information, SFT 2.0 is the better choice on every adequately sampled benchmark. For triage, neither model is acceptable — SFT 2.0 at 37.9% and SFT 3.0 at 40.2% are both below the untouched base model, and both must be kept out of any triage decision path.
Do not adopt SFT 3.0 over SFT 2.0 on the strength of the triage delta alone. The +2.3-point triage improvement is purchased with −1.1 to −3.8 on general QA, and triage is still not safe in either model.
Bias, Risks, and Limitations
Medical and clinical risks
- Fabricated clinical claims. The model will state a drug interaction, a contraindication, or a dose that does not exist, in fluent and confident prose. There is no factuality layer in the model itself.
- Systematic under-escalation. Documented above. The model downgrades true emergencies on the tested distribution.
- Over-triage of hopeless cases. Documented above. Causes alarm fatigue and consumes scarce clinical attention.
- No uncertainty signal. The model rarely says "I don't know." It is not trained to express calibrated uncertainty and has no mechanism that would make it do so.
- No citation behavior. The model does not reliably attribute claims to guidelines. Do not present its output as sourced.
- Guideline staleness. Training evidence includes the Nigeria Standard Treatment Guidelines 2022 and the 8th-edition Nigeria Essential Medicines List. Guidance changes. The model has no update mechanism.
Language and demographic risks
- English only, by design and by correction. 483 Twi/Akan rows were explicitly removed from the released dataset and replaced with English rows. The model will not reliably handle Nigerian Pidgin, and its behavior on Nigerian-language input is untested and unknown. A clinic serving non-English-speaking patients should not assume graceful degradation.
- Narrow geographic and health-system framing. The intended deployment context is Nigeria. Behavior in other health systems, or on conditions not represented in Nigerian and African guidance, is unmeasured.
- Data provenance skew. 46.9% of the training mix is MedMCQA and MedQA derivatives, i.e. Indian-origin medical entrance exam material. Real-world clinical language, particularly patient-facing Nigerian English, is thinly represented.
- Domain emphasis. The training mix is deliberately weighted toward maternal, child-health, and infectious-disease content. Performance elsewhere is weaker and unmeasured.
Sociotechnical risks
- Automation bias. A clinician shown a confident, concise, well-formatted answer is more likely to accept it than a diffuse one. The brevity that makes this model useful also makes it harder to challenge.
- Responsibility diffusion. Offloading triage to a tool on a shared clinic device can shift accountability without shifting liability.
- False reassurance. A short answer reads as complete. Absence of caveats is a formatting property here, not a safety property.
Recommendations
Users (both direct and downstream) should be made aware of the risks, biases and limitations above. In particular:
Do not deploy this model in any patient-facing triage pathway. The red-class failure is measured, reproducible, and unresolved.
If you evaluate it anyway, at minimum:
- Never let model output alone determine urgency. Keep a human in the loop for every case where the answer mentions danger signs, and keep paper IMCI chart booklets physically present as the reference.
- Run your own triage evaluation on your own population before any use. Do not reuse our numbers; the base rate and case mix in a real clinic will differ, and a 0.0 recall will not survive contact with a real case mix.
- Add a deterministic rule layer for emergencies. Hard-coded symptom-to-action rules should override the model. The model's contribution above that layer is marginal.
- Treat every output as a draft requiring verification against the current national guideline.
- Do not redistribute until the license is resolved. See License.
- Log and review every disagreement between the model and the clinician. That log is the fastest path to knowing whether this model is safe in your setting.
How to Get Started with the Model
This is a GGUF file, not a Transformers checkpoint. It is loaded with
llama.cpp, LM Studio, Ollama, or any GGUF-compatible runtime — not with
transformers.
# Download
huggingface-cli download Laptopllm/medLLM_V1_SFT_2.0_GGUF --local-dir ./medllm
# Run interactively
llama-cli \
-m ./medllm/medLLM_V1_SFT_2.0.Q4_K_M.gguf \
--ctx-size 4096 \
--threads 8 \
-p "<|im_start|>system
Give clear and concise medical information. Ask only necessary questions. Give urgent action first when needed.<|im_end|>
<|im_start|>user
My 2-year-old has had fever and vomiting since last night. What should I do?<|im_end|>
<|im_start|>assistant
"
Serve it as an OpenAI-compatible local endpoint:
llama-server -m ./path/medLLM_V1_SFT_2.0.Q4_K_M.gguf --ctx-size 4096 --port 8080
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="local",
messages=[
{"role": "system", "content": "Give clear and concise medical information. "
"Ask only necessary questions. Give urgent action first when needed."},
{"role": "user", "content": "Can I take ibuprofen 400mg while pregnant?"},
],
temperature=0,
max_tokens=256,
)
print(resp.choices[0].message.content)
LM Studio: search Laptopllm, load medLLM_V1_SFT_2.0, use the built-in chat
view. LM Studio applies the Qwen3.5 chat template automatically.
Prompting notes
- Use the ChatML format above. The model was trained with
tokenizer.apply_chat_template(...), which emits ChatML. PlainSystem:/User:/Assistant:formatting was also evaluated and performs similarly on triage, but ChatML is the training format and is the safer default. - Set
temperature=0for anything safety-relevant or reproducible. - Do not enable thinking/reasoning mode. This model was trained with zero
<think>tokens.enable_thinking=Trueproduces degraded output and, in prior evaluations, burned large amounts of generation budget for no gain. - Keep context at 4096. Training used
max_seq_length=1024for SFT and 4096 for CPT. Longer contexts are unlikely to help and may degrade the short-answer behavior the model was trained for. - Expect short answers. If you need a long response, you want a different model.
Training Details
Three stages, all on Modal NVIDIA L40S GPUs using Unsloth with LoRA. All
stages trained adapters in 4-bit and merged to 16-bit. Seed 3407 throughout.
Stage 1 — Continual Pre-Training (medical knowledge)
Domain-injected continued pre-training of Qwen/Qwen3.5-4B-Base on a mixed corpus of
medical text, African health guidance, and general replay data.
| Objective | Standard causal LM loss |
| Corpus | general/ (fineweb-edu), medical/ (African medical, global medical, patient education, safety/triage), reasoning/ (codeparrot, finemath, openr1-math) |
| Corpus size | ~717M curated tokens across 587,054 sources (unverified — see below) |
| LoRA | r=64, α=128, dropout 0.05, on q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Learning rate | 2e-5, cosine, warmup_steps=2500 |
| Epochs | 1 |
| Batch | 4 per device × 2 grad accum = 8 effective |
| Max sequence length | 4096, packing enabled |
| Optimizer | adamw_8bit |
| Precision | bf16 |
| Early stopping | patience 15 on eval_all_loss, eval every 2000 steps |
Corpus size caveat. The 717M / 587,054 figures appear only in project documentation and are not reproducible from any manifest, dataset card, or training log in the repository. A second project document states 502M. The replay fraction is not pinned in code — domains are loaded at whatever size they exist on disk, and the only held-out split is a ~120-row proportional sample used for early stopping. Treat the CPT corpus composition as under-documented. The published domain-mix ratios cannot be reconstructed from the repository.
Output: Laptopllm/medLLM_v1_cpt_16bit.
Stage 2 — Supervised Fine-Tuning (response behavior)
SFT started from the merged CPT checkpoint, not from the original base model.
| Init | Laptopllm/medLLM_v1_cpt_16bit (16-bit merge) |
| Objective | Causal LM on assistant tokens only; system and user tokens masked |
| LoRA | r=32, α=64, dropout 0.05, same 7 target modules |
| Learning rate | 5e-5, cosine, warmup_ratio=0.05 |
| Epochs | 3 (nominal; early stopping with patience 4, eval every 200 steps) |
| Batch | 8 per device × 4 grad accum = 32 effective |
| Max sequence length | 1024, packing disabled (padded, not concatenated) |
| Optimizer | adamw_8bit |
| Precision | bf16 |
| Early stopping | patience 4 on eval loss, load_best_model_at_end=True |
Loss masking was applied with:
train_on_responses_only("<|im_start|>user\n", "<|im_start|>assistant\n")
For clarify_then_answer conversations, both the clarification turn and the final
answer contribute to the loss. Packing was disabled for SFT specifically to prevent
one patient's conversation from becoming context for another, while remaining enabled
for CPT.
Output: Laptopllm/MEDLLM_V1_SFT_2.0 (merged 16-bit).
Stage 3 — Quantization (this artifact)
save_pretrained_merged(16-bit)
→ patch config.json: num_nextn_predict_layers=0, num_mtp_layers=0,
text_config.mtp_num_hidden_layers=0
→ convert_hf_to_gguf.py --outtype f16 (upstream ggml-org/llama.cpp)
→ llama-quantize q4_k_m
Unsloth's built-in save_pretrained_gguf could not be used. It triggers VLM
auto-detection and fails on assert self.opt_num_mtp_layers != 0 for Qwen3.5. The
upstream llama.cpp path was vendored and the MTP assertion in conversion/qwen.py
pre-patched at image build time. Only q4_k_m is produced; no other quantization
level exists.
Training Data
SFT dataset — 30,000 conversations, one JSONL schema:
(id, messages, interaction_style, answer_policy, evidence, source, family_id, split)
Split: 27,000 train / 1,500 dev / 1,500 internal test (90/5/5), assigned by
family_id so all versions of a case stay in one split. Only provider training
splits were used; provider validation and test rows were excluded.
| Source | Rows | Share |
|---|---|---|
| MedMCQA (train) | 5,002 | 16.7% |
| DDXPlus | 4,500 | 15.0% |
| MedMCQA (legacy) | 3,899 | 13.0% |
| MedQuAD | 3,385 | 11.3% |
| MedQA USMLE (legacy) | 3,030 | 10.1% |
| AfriHealth-QA (train) | 2,978 | 9.9% |
| MedQA US (train) | 2,127 | 7.1% |
| PubMedQA (legacy) | 1,499 | 5.0% |
| MedSafety-Improve (train) | 897 | 3.0% |
| ChatDoctor (legacy) | 704 | 2.3% |
| MEDEC-MS | 544 | 1.8% |
| Medical-O1 reasoning-SFT (legacy, final answer only) | 468 | 1.6% |
| MedicationQA | 467 | 1.6% |
| MedlinePlus (topics) | 411 | 1.4% |
| AfriMed-QA v2 (train) | 89 | 0.3% |
| Total | 30,000 | 100% |
English-only, by explicit correction. 483 Twi/Akan rows (435 train, 27 dev, 21 internal test), identified by Akan diacritics, were removed and replaced 1:1 with unused English rows from the same AfriHealth-QA parquet. Replacements were filtered for no chain-of-thought tags, no non-answers, no administrative language, answers ≤180 words, and no truncated sentence endings. Final Twi/Akan count: 0.
Preprocessing [optional]
Applied to the SFT corpus:
- ChatDoctor boilerplate stripping via opener and closer regex sets
- Length filtering: drop rows <40 or >2000 characters
- Intra-source deduplication on lowercased instruction
- Cross-source deduplication on normalized instruction
- Holdout-overlap removal between Medical-O1 reasoning rows and the verifiable holdout (the two families share ~33% of questions)
- Empty-field skips
- Re-render through
tokenizer.apply_chat_template(convo, add_generation_prompt=False)
Training Hyperparameters
Consolidated across all three stages:
| Parameter | CPT | SFT | Repair pilot |
|---|---|---|---|
| LoRA r / α | 64 / 128 | 32 / 64 | 16 / 32 |
| LoRA dropout | 0.05 | 0.05 | 0.05 |
| Target modules | q,k,v,o,gate,up,down proj | same | same |
| Learning rate | 2e-5 | 5e-5 | 1e-5 |
| Scheduler | cosine | cosine | cosine |
| Warmup | 2500 steps | 5% ratio | 5% ratio |
| Epochs | 1 | 3 | 1 |
| Effective batch | 8 | 32 | 32 |
| Max sequence length | 4096 | 1024 | 1024 |
| Packing | on | off | off |
| Optimizer | adamw_8bit |
adamw_8bit |
adamw_8bit |
| Training regime | bf16 mixed precision | bf16 mixed precision | bf16 mixed precision |
| Seed | 3407 | 3407 | 3407 |
| Early stopping | patience 15 | patience 4 | disabled (all ckpts kept) |
| Checkpoints | every 2000, limit 4 | every 200, limit 3 | 35 / 70 / 105 / 140 |
The repair pilot was run on top of SFT 2.0 with LoRA r=16, α=32, lr=1e-5, 1 epoch, effective batch 32, on 4,500 train / 500 dev rows (1,800 triage + 2,700 replay). It was not merged and not released — its red-recall gate failed.
Speeds, Sizes, Times [optional]
- Parameters: 4,205,751,296
- GGUF Q4_K_M artifact: ~2.7 GB
- MMLU evaluation of the CPT checkpoint: 765.4 s total on 2× Tesla T4, bf16,
zero-shot,
lm-eval 0.4.12 - Modal L40S training. Exact training wall-clock hours were not recorded; see Environmental Impact.
Framework versions
- Unsloth
2026.8.4(saved in model config);>=2026.7.20pinned in SFT scripts - Transformers
5.5.0 - PEFT
0.20.0 - TRL
0.24.0 - PyTorch
2.10.0 - Datasets
4.3.0, Tokenizers0.22.2 - lm-eval
0.4.12(MMLU evaluation only) - llama.cpp: upstream
ggml-org/llama.cpp, vendored and patched - No DeepSpeed, FSDP, or ZeRO configuration was used.
Evaluation
All numbers below were produced with llama.cpp inference, temperature 0, fixed
seed, thinking disabled, and matched quantization across all compared checkpoints.
Testing Data, Factors & Metrics
Testing Data
| Benchmark | Cases | Note |
|---|---|---|
| MedQA USMLE 4-option (test) | 1,273 | Untouched test split; train rows were used in SFT |
| MedMCQA (validation) | 4,183 | Untouched validation split; train rows were used in SFT |
| AfriMed-QA | 3,910 | Pan-African; row-level train/test filter required |
| NigeriaMedQA | 844 / 909 | Physician-reviewed, Nigerian guidelines. Count discrepancy unresolved. |
| MedExpQA English (test) | 125 | Not an SFT source family — the cleanest transfer check |
| Triage (internal dev) | 200 | 50 per class, counterfactual families kept atomic |
| Triage (TRIAGE benchmark) | 87 | Mass-casualty START/jumpSTART-style only |
| ACI-Bench | 8 | Far too small for the reported deltas |
| MTS-Dialog | 15 | Far too small for the reported deltas |
| Safety benchmarks | 0 | HealthBench, PatientSafetyBench, MedSafetyBench, LiveQA not run |
Factors
Results are disaggregated by per-class precision/recall/F1 for triage, and by benchmark for QA. African relevance is reported as NigeriaMedQA and AfriMed-QA. No demographic disaggregation exists — there is no evaluation by patient age, sex, region, language, or care setting, and none is possible with the current data.
Metrics
- MCQ: accuracy with Wilson 95% CI
- Triage: macro-F1 as primary; critical under-triage (red recall) reported first; per-class precision/recall/F1; unparsed-output rate
- Zeros matter: a 0.0 recall is reported as a gate failure, not averaged away against a good macro-F1
- Regression gating: no benchmark may drop more than 2 points, and general-average may not drop more than 1 point, for a checkpoint to be accepted
Results
Base → CPT → SFT 2.0, identical protocol:
| Benchmark | Cases | Base | CPT | SFT 2.0 | SFT 2.0 vs base |
|---|---|---|---|---|---|
| MedQA USMLE 4-opt (test) | 1,273 | 62.8 | 56.9 | 64.0 | +1.2 |
| MedMCQA (validation) | 4,183 | 54.1 | 51.4 | 58.4 | +4.3 |
| AfriMed-QA | 3,910 | 59.1 | 57.6 | 62.4 | +3.3 |
| MedExpQA English (test) | 125 | 56.0 | 55.2 | 62.4 | +6.4 |
| NigeriaMedQA | 844 | 79.3 | 78.2 | 79.1 | −0.2 |
| Triage | 87 | 43.7 | 42.5 | 37.9 | −5.8 |
| ACI-Bench | 8 | 23.4 | 18.7 | 13.6 | −9.8 |
| MTS-Dialog | 15 | 18.8 | 7.8 | 8.6 | −10.2 |
This model beats the base on four medical QA benchmarks and loses on four others, two of them heavily. The two large losses are in clinical note generation (ACI-Bench, MTS-Dialog) and the triage loss is safety-relevant.
Sample-size warning: ACI-Bench (8), MTS-Dialog (15), and Triage (87) are too small for the point differences shown to be meaningful. Only MedQA, MedMCQA, and AfriMed-QA are adequately sampled.
MMLU medical subsets — CPT-stage result only
Measured on the CPT checkpoint (Laptopllm/medLLM_v1_cpt_16bit), zero-shot,
lm-eval 0.4.12, bf16, 2× Tesla T4, 4,205,751,296 parameters:
| Task | n | Accuracy |
|---|---|---|
| MMLU college biology | 144 | 88.89% |
| MMLU professional medicine | 272 | 81.25% |
| MMLU medical genetics | 100 | 84.00% |
| MMLU clinical knowledge | 265 | 80.75% |
| MMLU college medicine | 173 | 76.30% |
| MMLU anatomy | 135 | 74.07% |
These are not SFT 2.0 results. They are reported because they are the only high-quality medical-knowledge measurements available for this model family. The final model averaged ~80.2% across MMLU medical subsets, below the plotted base and CPT averages. MMLU is a regression result here, not a headline improvement.
Triage — the critical result
SFT 2.0 on the 200-case balanced internal triage dev set:
| Metric | Value |
|---|---|
| Accuracy | 44.5% |
| Macro-F1 | 43.42% |
| Unparsed outputs | 0 |
| Red precision / recall / F1 | 0.0 / 0.0 / 0.0 |
| Yellow recall | 36.0% |
| Green recall | 90.0% |
| Black recall | 52.0% |
Full per-class and confusion-matrix detail is in Known Critical Limitations.
Repair stage, gated evaluation against the SFT 2.0 baseline:
| Config | dev200 acc | macro-F1 | Red recall | Y/G/B recall | MedQA | MedMCQA | AfriMedQA | NigMedQA | MedExpQA | Gates |
|---|---|---|---|---|---|---|---|---|---|---|
| baseline (SFT 2.0) | 44.5 | 43.42 | 0.0 | 36 / 90 / 52 | 66.67 | 58.67 | 65.33 | 75.33 | 62.4 | — |
| ckpt-70 | 61.0 | 52.69 | 0.0 | 44 / 100 / 100 | 66.00 | 59.33 | 65.33 | 75.33 | 63.2 | 4/5 |
| ckpt-105 | 62.5 | 54.10 | 0.0 | 50 / 100 / 100 | 66.67 | 58.67 | 65.33 | 75.33 | 64.8 | 4/5 |
| ckpt-140 | 64.5 | 56.02 | 0.0 | 58 / 100 / 100 | 66.00 | 58.67 | 65.33 | 75.33 | 62.4 | 4/5 |
The repair gate required: Recall_red ≥ 85%, Y/G/B recall non-decreasing,
Δmacro-F1 > 0, Δgeneral-average ≥ −1%, no benchmark down >2%. All three checkpoints
passed the four quality gates and failed the red-recall gate. The repair model was
not merged or released.
TRIAGE benchmark (87 cases), under both ChatML and a plain
System:/User:/Assistant: harness:
| Config | ChatML | Plain harness |
|---|---|---|
| SFT 2.0 baseline | 36.78 | 40.23 |
| ckpt-140 | 40.23 | 41.38 |
Summary
MEDLLM V1 SFT 2.0 is a fast, small, fully offline English-language medical assistant that beats its base model on three adequately-sampled medical QA benchmarks and is meaningfully worse on clinical note generation. It cannot perform emergency triage, and a targeted repair stage failed to fix that. It is suitable for patient education and general information in an offline setting. It is not suitable for triage, diagnosis, or any decision where a wrong answer harms a patient.
For how this model compares to the later SFT 3.0 release, see Comparison to SFT 3.0.
Model Examination [optional]
No interpretability work has been done on this model. No attention analysis, no probing, no representation study, no mechanistic interpretability.
Two observations from the evaluation are worth recording as hypotheses for future work, stated as hypotheses rather than findings:
- The
redclass appears unrepresentable rather than merely under-weighted. The confusion matrix shows all 24redpredictions landing onblackgolds, and zero onredgolds. That pattern is more consistent with a mislabeled or semantically invertedreddefinition in the training data than with a simple class-prior problem — if it were a prior problem,redpredictions would scatter across classes rather than concentrating on one. This should be checked first before spending more compute on repair training. - The black→red confusion is a potential data-quality signal. If the label definition maps "expected to die" onto "emergency" somewhere in the authored set, both symptoms follow from one bug. Auditing the label definitions would resolve both at once.
Environmental Impact
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
- Hardware Type: NVIDIA L40S (48 GB), rented via Modal
- Hours used: [Not recorded — training wall-clock hours were not logged and cannot be recovered from the repository]
- Cloud Provider: Modal
- Compute Region: [Not recorded]
- Carbon Emitted: [Not estimated — the input above is missing]
Context on why this number is not filled in: rather than fabricate a figure, note
that all training was on rented L40S instances rather than local hardware, and that
GPU access — not carbon — was the binding constraint on this project. An early
learning-rate pilot diverged and consumed a paid run before the schedule was
restarted conservatively. Recorded evaluation cost was also material: a matched
10,422-question llama.cpp sweep plus the Medical MMLU comparison were batched
specifically to avoid repeated paid evaluations.
A rough estimate is possible: multiply Modal L40S-hours by a published L40S emissions-per-hour figure, including both training and the failed/repeated runs described in the project documentation. This has not been done, and any number published without the actual hour count would be a guess.
Technical Specifications [optional]
Model Architecture and Objective
Derived from Qwen3.5-4B-Base, architecture Qwen3_5ForConditionalGeneration.
| Property | Value |
|---|---|
| Total parameters | 4,205,751,296 |
| Architecture | Causal decoder, hybrid attention |
num_hidden_layers |
32 |
| Layer pattern | 3× linear_attention + 1× full_attention, repeating |
full_attention_interval |
4 |
hidden_size |
2,560 |
intermediate_size |
9,216 |
num_attention_heads |
16 |
num_key_value_heads |
4 (GQA) |
head_dim |
256 |
linear_num_key_heads / linear_num_value_heads |
16 / 32 |
linear_key_head_dim / linear_value_head_dim |
128 / 128 |
linear_conv_kernel_dim |
4 |
attn_output_gate |
true |
partial_rotary_factor |
0.25 |
rope_theta |
10,000,000 |
vocab_size |
248,320 |
tie_word_embeddings |
true |
max_position_embeddings |
262,144 |
| Activation | silu |
eos_token_id |
248,044 |
| Objective | Causal language modeling |
MTP head removed. The base config declares mtp_num_hidden_layers: 1 and 33
blocks, but the merged fine-tune contains only blk.0–blk.31. This export sets
num_nextn_predict_layers=0, num_mtp_layers=0, and
text_config.mtp_num_hidden_layers=0. The multi-token-prediction head is a decode
throughput optimization, not a quality component, so removing it does not affect
output quality. Without this patch LM Studio fails with a
blk.32.attn_norm.weight error and upstream convert_hf_to_gguf.py asserts on
opt_num_mtp_layers.
Vision tower. The config.json retains a vision_config block (depth 24,
hidden_size 1024, patch size 16, spatial merge 2). It is inherited from the base
architecture and is not trained, not used, and not exercised by this
model. This is a text-only model. Ignore the vision block.
Compute Infrastructure
- Training: Modal, NVIDIA L40S (48 GB), Unsloth fast path
- Adapter precision: 4-bit quantized load, LoRA adapters in bf16 compute
- Merge: 16-bit (
save_pretrained_merged,merged_16bit) - Inference:
llama.cppon CPU; no GPU required - Context: 4096 recommended (262,144 theoretical max; untrained at long context)
- DeepSpeed / FSDP / ZeRO: not used
Hardware
- Training: 1× NVIDIA L40S 48 GB per run
- Evaluation: 2× Tesla T4 for MMLU; CPU for all
llama.cppbenchmark sweeps - Deployment target: 8 GB consumer laptop, CPU-only
Software
| Component | Version |
|---|---|
| Unsloth | 2026.8.4 (saved in config); >=2026.7.20 pinned in scripts |
| Transformers | 5.5.0 |
| PEFT | 0.20.0 |
| TRL | 0.24.0 |
| PyTorch | 2.10.0 |
| Datasets | 4.3.0 |
| Tokenizers | 0.22.2 |
| lm-eval | 0.4.12 |
| llama.cpp | upstream ggml-org/llama.cpp, vendored, conversion/qwen.py patched |
| Python | 3.11 (training), 3.12.13 (evaluation) |
License
This model has no declared license and must not be redistributed.
The repository contains exactly one LICENSE file, docker/LICENSE, which is MIT
and covers the Docker tooling only. It does not cover the model weights, the
training data, or the datasets. No MEDLLM model or dataset declares a license in its
model card, dataset card, or release manifest.
This is unresolved and must be settled before publication or distribution. It
directly conflicts with the Qwen3.5-4B-Base upstream terms, which a derivative
must respect, and with several training-data sources whose own terms restrict use —
MedSafetyBench is explicitly research-only, and AfriMed-QA's canonical
repository and public mirror show conflicting terms.
The correct license is a decision for the model owner, informed by the upstream
base-model terms and by the licensing status of every training source. It is not a
factual matter that can be inferred from the repository. The frontmatter declares
license: other as a placeholder to block automated tooling from assuming a
permissive default.
Note also that MedSafetyBench data is research-only and TRIAGE declares no
license. Shipping a model trained on those sources requires resolving each.
Citation [optional]
No paper or blog post introducing this model has been published.
BibTeX: [More Information Needed]
APA: [More Information Needed]
If you use this model, please cite the base model
(Qwen/Qwen3.5-4B-Base) and the
llama.cpp project.
Glossary [optional]
| Term | Meaning |
|---|---|
| CPT | Continual Pre-Training. Continued training of an existing model on additional domain data. Here: adding medical knowledge to the Qwen3.5-4B base. |
| SFT | Supervised Fine-Tuning. Training on curated input/output demonstrations to control response style. |
| LoRA | Low-Rank Adaptation. Trains a small low-rank matrix alongside frozen weights instead of full fine-tuning. |
| assistant-only loss | Loss computed only on assistant tokens; system and user tokens are masked out. Prevents the model learning to echo the prompt. |
| packing | Concatenating multiple short examples into one fixed-length sequence. Improves throughput for pretraining; disabled for SFT here so one patient's conversation cannot become context for another. |
| MTP | Multi-Token Prediction. A head that predicts several future tokens per step to speed up decoding. Removed in this export. |
| Q4_K_M | A 4-bit GGUF quantization using k-quant mixed precision. The quality/size tradeoff used for the laptop build. |
| GGUF | The file format used by llama.cpp and LM Studio. Not loadable by transformers. |
| red / yellow / green / black | Four-label triage scheme. Red = immediate emergency, yellow = urgent but not emergent, green = routine, black = expected to die / palliative or expectant. |
| Under-triage | Assigning a lower urgency label than clinically warranted. The primary safety failure mode measured here. |
| Over-triage | Assigning a higher urgency label than warranted. Causes resource and attention misallocation. |
| Macro-F1 | Mean F1 across all classes, weighting each class equally. Can mask a total failure on one class — hence red recall is reported separately. |
| Family ID | A stable identifier grouping all versions of the same underlying case, used to prevent leakage across train/dev/test splits. |
| IMCI | Integrated Management of Childhood Illness. A WHO clinical guideline for child health, used as an evidence source. |
| NEML | Nigeria Essential Medicines List. |
| NCDC | Nigeria Centre for Disease Control. |
| STP / NTBLCP | National Tuberculosis and Leprosy Control Programme. |
More Information [optional]
Intended design vs. delivered artifact
The project's stated design goal was a locally-grounded, Nigerian-context medical assistant. Two documented gaps separate that goal from this artifact:
- The Nigerian official-guidance rows are not in the released dataset. The specification called for 6,070 such rows plus 8,460 controlled authored cases. Neither appears. See Known Critical Limitations.
- The CPT stage regressed general capability rather than adding knowledge cleanly. See Known Critical Limitations.
Both are recoverable. The next steps identified in project documentation are: resolve
the triage label definition before further repair training; select and merge a
triage-repair checkpoint once the red gate can actually be passed; add retrieval over
NCDC/NEML/IMCI guidance with evidence.url + locator citations so grounding comes
from retrieval rather than weights; expand Nigerian English and Pidgin coverage with
a language: pcm evaluation slice; and run a prospective clinic pilot against paper
IMCI booklets measuring latency, triage concordance, and clinician trust on-device.
Repository and artifacts
- Training scripts:
qwen-cpt.py,sft_modal.py,sft_repair_modal.py - Data preparation:
prepare_data.py,prepare_repair_data.py,verify_data.py - Evaluation:
eval_diagnostic_modal.py,eval_repair_modal.py,generate_eval_charts.py - Export:
finetune.py,covert_gguf.py(superseded) - Schema:
unified-medical-sft.schema.json - Scoreboards:
repair_eval_results/REPAIR-EVAL-SCOREBOARD.md,MEDLLM-EXACT-EVALS-MODAL/.../reference/previous-4090-SCOREBOARD.md - Model cards:
MODEL_CARD_SFT_2.0_GGUF.md(original), this file (revision with SFT 3.0 comparison)
All model repos were created with private=True.
Model Card Authors [optional]
Arinde David (Laptopllm). [More Information Needed — confirm contributors and
contributions before publication]
Model Card Contact [optional]
[More Information Needed]
This card was written directly from repository artifacts: training scripts,
config.json, the SFT build manifest, the triage repair scoreboard, and internal
evaluation scoreboards. Every number above is traceable to one of those files.
Discrepancies found between sources are flagged inline rather than resolved
silently. The SFT 3.0 comparison draws on the same evaluation harness and the
37,000-row triage-augmented build manifest.
- Downloads last month
- 36
4-bit
Model tree for Laptopllm/MedAssist
Base model
Laptopllm/medLLM_v1_cpt_16bit