KneeVision-models / README.md
Darshan13's picture
Upload README.md with huggingface_hub
f89f6b3 verified
|
Raw History Blame Contribute Delete
3.34 kB
metadata
license: mit
tags:
  - medical-imaging
  - knee-osteoarthritis
  - kellgren-lawrence
  - x-ray
  - multimodal
  - pytorch

KneeVision++ — Knee Osteoarthritis Severity Grading Checkpoints

Inference checkpoints from KneeVision++, an explainable multimodal pipeline for grading knee osteoarthritis severity (Kellgren-Lawrence 0–4) from X-rays and/or structured clinical/radiographic text.

Live demo: Darshan13/kneevision-demo

Files

File Model Task Test Quadratic Kappa
best_densenet121.pt DenseNet121 5-class KL grading (image) 0.775
best_efficientnet-b4.pt EfficientNet-B4 5-class KL grading (image) 0.640
best_convnext_small_binary.pt ConvNeXt-Small binary OA screening (image) 0.757
best_densenet121_binary.pt DenseNet121 binary OA screening (image) 0.719
best_efficientnet-b4_binary.pt EfficientNet-B4 binary OA screening (image) 0.658
best_clinical.pt BioClinicalBERT 5-class KL grading (clinical text) 0.953
best_fusion.pt CNN + BioClinicalBERT fusion 5-class KL grading (multimodal) 0.959

Each best_*.pt has a matching best_*.json sidecar with {model_name, num_classes, binary, ordinal, image_size, best_kappa} metadata used by the loader.

Training data

  • Images: Kaggle Knee Osteoarthritis Dataset — 8,260 X-rays, KL grades 0–4
  • Clinical text: composed from real Osteoarthritis Initiative (OAI) fields — demographics, WOMAC pain/stiffness/function, and per-compartment OARSI radiographic grades (joint space narrowing, osteophytes, sclerosis, attrition) — not the KL label itself

Test-set results (1,656 held-out X-rays)

Modality Accuracy Quadratic Kappa
Image only (DenseNet121) 60.3% 0.792
Clinical Text only (BioClinicalBERT) 87.4% 0.953
Multimodal Fusion 88.9% 0.959

Caveat: the clinical-text branch is fed real per-compartment OARSI fields that mechanically define the KL grade, not free-text patient-reported symptoms — so this is closer to "decode KL's own defining components written as prose" than "infer severity from symptoms alone." Fusion only marginally beats text-alone (+1.5pp accuracy) because the text branch already carries near-complete radiographic signal. Full discussion in the project README.

Usage

from huggingface_hub import hf_hub_download
import torch

path = hf_hub_download(repo_id="Darshan13/KneeVision-models", filename="best_densenet121.pt")
state = torch.load(path, map_location="cpu", weights_only=False)

See src/kneevision/models/image_model.py (load_trained_model) for the full loading logic (backbone reconstruction, ordinal-head detection, multitask auxiliary-head detection).

Disclaimer

Research/educational checkpoints only. Not a medical device, not validated for clinical use, not intended to inform real diagnostic or treatment decisions.