Instructions to use drdai/diagnify-14b-v0.2-cot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use drdai/diagnify-14b-v0.2-cot with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="drdai/diagnify-14b-v0.2-cot") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("drdai/diagnify-14b-v0.2-cot", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use drdai/diagnify-14b-v0.2-cot with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "drdai/diagnify-14b-v0.2-cot" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "drdai/diagnify-14b-v0.2-cot", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/drdai/diagnify-14b-v0.2-cot
- SGLang
How to use drdai/diagnify-14b-v0.2-cot with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "drdai/diagnify-14b-v0.2-cot" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "drdai/diagnify-14b-v0.2-cot", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "drdai/diagnify-14b-v0.2-cot" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "drdai/diagnify-14b-v0.2-cot", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use drdai/diagnify-14b-v0.2-cot with Docker Model Runner:
docker model run hf.co/drdai/diagnify-14b-v0.2-cot
DIAGNIFY — 14B · v0.2-CoT
Expert-mode diagnostic reasoning. Retrieval-first. Severity-weighted.
DIAGNIFY is a 14B-parameter clinical reasoning system developed by drdai. It is designed to retrieve evidence from a curated, structured medical knowledge base, reason through body systems using an explicit chain-of-thought (CoT) framework, and produce a structured differential diagnosis ranked by likelihood × clinical severity.
The system is designed to surface life-threatening conditions before committing to a final diagnosis. Retrieval grounding is intended to support evidence traceability; it does not, by itself, guarantee factual accuracy or prove that an explanation faithfully represents the model's internal computation.
The curated knowledge base is reported by the developer to be independent of all evaluation cases.
At a Glance
| Specification | Value |
|---|---|
| Model | drdai/diagnify-14b-v0.2-cot |
| Parameters | 14B |
| System architecture | Retrieval-Augmented Generation (RAG) + CoT |
| Knowledge base | Curated and structured; reported to be disjoint from evaluation cases |
| License | Apache-2.0 |
| Status | Research artifact; not approved for clinical use |
How the System Works
CASE VIGNETTE
|
v
[RETRIEVE] -> Curated, structured knowledge-base lookup
|
v
[REASON] -> System-by-system reasoning framework
|
v
[FLAG] -> Can't-miss and life-threatening conditions
|
v
[RANK] -> Likelihood x clinical severity weighting
|
v
[COMMIT] -> Final diagnosis + evidence-linked rationale
The intended workflow exposes retrieved evidence, a structured diagnostic rationale, safety flags, and the final diagnostic ranking for research review. Source support must be checked rather than assumed from the presence of a reasoning trace.
Capabilities
- Structured differentials across multi-system presentations.
- Severity-prioritized flagging to elevate dangerous conditions for consideration, including conditions with a low base rate.
- Rare and atypical case coverage supported by retrieval grounding.
- Evidence-linked diagnostic rationales for research review and validation.
Evaluation
Warning: Single preliminary internal benchmark. Not independently replicated. Not clinical validation.
All results below are developer-reported and should be interpreted within the limitations of the internal evaluation.
Dataset and Grading
The reported dataset comprises approximately 2,500+ real, de-identified patient cases contributed by Australian academic GPs. The developer reports existing consent and de-identification of the cases.
Cases were selected for diagnostic difficulty, including rare, multi-system, and atypical presentations. Grading was reported to have been performed independently by practising RACGP fellow physicians.
Reported Results
| System | Reported diagnostic-label accuracy |
|---|---|
| DIAGNIFY (RAG + CoT) | ~100%* |
| GPT-5.5 | 83% |
| Claude Opus 4.8 | 83% |
| Gemini 3.1 Pro | 75% |
| Grok 4.20 | 58% |
| DeepSeek V4 | 42% |
Secondary Metrics
The following findings carry the same preliminary internal-benchmark caveats:
- Reported high accuracy on ultra-rare and multi-system subsets.
- Reported low unnecessary-test recommendation rate.
- Reported grader agreement: Cohen's κ ≈ 0.94.
Important Caveat About the ~100% Result
*The reported ~100% accuracy should not be treated as established diagnostic performance.
This figure reflects the conditions of this specific internal benchmark, including case selection, grading criteria, and knowledge-base overlap controls. It does not establish near-perfect performance on new cases, ambiguous presentations, or real clinical workflows.
The result requires independent, prospective replication using held-out cases, documented leakage controls, and blinded external grading.
What Was Measured
Diagnostic-label accuracy only. The evaluation did not measure patient outcomes or performance within a live clinical workflow.
Safety, Boundaries, and Regulatory Status
NOT FOR CLINICAL USE.
This release is intended for research and experimentation only. It must not be relied upon to make, delay, or replace clinical decisions.
| Boundary | Declared status or intended-use restriction |
|---|---|
| Regulatory clearance or approval | Not cleared. Not approved. Not submitted. |
| Clinical decision-making | Outside intended use. Must not be relied upon for diagnosis or treatment decisions. |
| Patient-facing deployment | Outside intended use. Research and experimentation only. |
| Human oversight | All medical outputs require verification by qualified clinicians. |
Before any proposed clinical deployment: Appropriate clinical validation, safety evaluation, and regulatory review would be required. This research benchmark does not establish readiness for clinical use.
Limitations
- Static vignettes only. The reported evaluation did not assess live history-taking, physical examination, or dynamic test interpretation.
- Selection bias. Cases lean toward documented, relatively classic rare presentations and may not reflect ambiguous primary-care or emergency presentations.
- Interface confounding. System comparisons may reflect differences in interaction design, retrieval access, or evaluation conditions rather than reasoning capability alone.
- Label accuracy is not outcome improvement. Correctly naming a diagnosis does not establish clinical benefit or harm reduction.
- Retrieval is not a guarantee of correctness. Retrieved material and generated rationales require verification.
- Limited reproducibility information. Exact case counts, per-case results, model snapshots, prompts, grading criteria, and retrieval configurations are not specified in this card.
These results constitute a preliminary research benchmark, not clinical or regulatory validation.
Intended Use
- Research on retrieval-grounded clinical reasoning.
- Structured reasoning frameworks for differential diagnosis.
- Benchmarking diagnostic AI architectures.
- Non-clinical decision-support prototyping with qualified human oversight.
Explicitly out of scope: Autonomous diagnosis, patient triage, treatment decisions, and clinical deployment of this research release.
Quick Start
Scope of this example: This code loads the language-model component and accepts externally retrieved context. It does not implement the curated knowledge base, retrieval pipeline, or full DIAGNIFY orchestration. It requires access to the model repository and compatible model artifacts.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "drdai/diagnify-14b-v0.2-cot"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
case = "<de-identified clinical vignette>"
# Retrieve relevant evidence using your knowledge-base pipeline first.
# Replace this placeholder with actual source-labelled excerpts.
retrieved_context = "<retrieved evidence with source identifiers>"
prompt = f"""You are DIAGNIFY, a research-only clinical reasoning system.
Use the supplied retrieved evidence to support your analysis.
Do not claim to have retrieved sources that were not supplied.
Identify evidence gaps and uncertainty.
Work through relevant body systems.
Flag potentially life-threatening conditions.
Distinguish diagnostic likelihood from clinical urgency.
Provide a structured differential, an evidence-linked rationale,
and a final diagnostic hypothesis for qualified clinician review.
RETRIEVED EVIDENCE:
{retrieved_context}
CASE:
{case}
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=False,
)
# Decode only the generated response, excluding the input prompt.
generated_tokens = outputs[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated_tokens, skip_special_tokens=True))
Important: The full DIAGNIFY system includes the curated knowledge base, retrieval pipeline, severity-ranking logic, and orchestration layer. Loading the model weights alone does not reproduce the complete RAG system or the reported benchmark results.