File size: 3,336 Bytes
f89f6b3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 | ---
license: mit
tags:
- medical-imaging
- knee-osteoarthritis
- kellgren-lawrence
- x-ray
- multimodal
- pytorch
---
# KneeVision++ — Knee Osteoarthritis Severity Grading Checkpoints
Inference checkpoints from **[KneeVision++](https://github.com/Darshan-T-P/KneeVision)**, an
explainable multimodal pipeline for grading knee osteoarthritis severity (Kellgren-Lawrence 0–4)
from X-rays and/or structured clinical/radiographic text.
Live demo: [Darshan13/kneevision-demo](https://huggingface.co/spaces/Darshan13/kneevision-demo)
## Files
| File | Model | Task | Test Quadratic Kappa |
|---|---|---|--:|
| `best_densenet121.pt` | DenseNet121 | 5-class KL grading (image) | 0.775 |
| `best_efficientnet-b4.pt` | EfficientNet-B4 | 5-class KL grading (image) | 0.640 |
| `best_convnext_small_binary.pt` | ConvNeXt-Small | binary OA screening (image) | 0.757 |
| `best_densenet121_binary.pt` | DenseNet121 | binary OA screening (image) | 0.719 |
| `best_efficientnet-b4_binary.pt` | EfficientNet-B4 | binary OA screening (image) | 0.658 |
| `best_clinical.pt` | BioClinicalBERT | 5-class KL grading (clinical text) | 0.953 |
| `best_fusion.pt` | CNN + BioClinicalBERT fusion | 5-class KL grading (multimodal) | 0.959 |
Each `best_*.pt` has a matching `best_*.json` sidecar with `{model_name, num_classes, binary,
ordinal, image_size, best_kappa}` metadata used by the loader.
## Training data
- **Images:** [Kaggle Knee Osteoarthritis Dataset](https://www.kaggle.com/datasets/shashwatwork/knee-osteoarthritis-dataset-with-severity) — 8,260 X-rays, KL grades 0–4
- **Clinical text:** composed from real [Osteoarthritis Initiative (OAI)](https://nda.nih.gov/oai) fields — demographics, WOMAC pain/stiffness/function, and per-compartment OARSI radiographic grades (joint space narrowing, osteophytes, sclerosis, attrition) — not the KL label itself
## Test-set results (1,656 held-out X-rays)
| Modality | Accuracy | Quadratic Kappa |
|---|--:|--:|
| Image only (DenseNet121) | 60.3% | 0.792 |
| Clinical Text only (BioClinicalBERT) | 87.4% | 0.953 |
| Multimodal Fusion | 88.9% | 0.959 |
**Caveat:** the clinical-text branch is fed real per-compartment OARSI fields that mechanically
define the KL grade, not free-text patient-reported symptoms — so this is closer to "decode KL's
own defining components written as prose" than "infer severity from symptoms alone." Fusion only
marginally beats text-alone (+1.5pp accuracy) because the text branch already carries near-complete
radiographic signal. Full discussion in the [project README](https://github.com/Darshan-T-P/KneeVision#phase-5--multimodal-fusion-pipeline-).
## Usage
```python
from huggingface_hub import hf_hub_download
import torch
path = hf_hub_download(repo_id="Darshan13/KneeVision-models", filename="best_densenet121.pt")
state = torch.load(path, map_location="cpu", weights_only=False)
```
See [`src/kneevision/models/image_model.py`](https://github.com/Darshan-T-P/KneeVision/blob/main/src/kneevision/models/image_model.py)
(`load_trained_model`) for the full loading logic (backbone reconstruction, ordinal-head detection,
multitask auxiliary-head detection).
## Disclaimer
Research/educational checkpoints only. Not a medical device, not validated for clinical use, not
intended to inform real diagnostic or treatment decisions.
|