File size: 3,336 Bytes
f89f6b3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
---
license: mit
tags:
  - medical-imaging
  - knee-osteoarthritis
  - kellgren-lawrence
  - x-ray
  - multimodal
  - pytorch
---

# KneeVision++ — Knee Osteoarthritis Severity Grading Checkpoints

Inference checkpoints from **[KneeVision++](https://github.com/Darshan-T-P/KneeVision)**, an
explainable multimodal pipeline for grading knee osteoarthritis severity (Kellgren-Lawrence 0–4)
from X-rays and/or structured clinical/radiographic text.

Live demo: [Darshan13/kneevision-demo](https://huggingface.co/spaces/Darshan13/kneevision-demo)

## Files

| File | Model | Task | Test Quadratic Kappa |
|---|---|---|--:|
| `best_densenet121.pt` | DenseNet121 | 5-class KL grading (image) | 0.775 |
| `best_efficientnet-b4.pt` | EfficientNet-B4 | 5-class KL grading (image) | 0.640 |
| `best_convnext_small_binary.pt` | ConvNeXt-Small | binary OA screening (image) | 0.757 |
| `best_densenet121_binary.pt` | DenseNet121 | binary OA screening (image) | 0.719 |
| `best_efficientnet-b4_binary.pt` | EfficientNet-B4 | binary OA screening (image) | 0.658 |
| `best_clinical.pt` | BioClinicalBERT | 5-class KL grading (clinical text) | 0.953 |
| `best_fusion.pt` | CNN + BioClinicalBERT fusion | 5-class KL grading (multimodal) | 0.959 |

Each `best_*.pt` has a matching `best_*.json` sidecar with `{model_name, num_classes, binary,
ordinal, image_size, best_kappa}` metadata used by the loader.

## Training data

- **Images:** [Kaggle Knee Osteoarthritis Dataset](https://www.kaggle.com/datasets/shashwatwork/knee-osteoarthritis-dataset-with-severity) — 8,260 X-rays, KL grades 0–4
- **Clinical text:** composed from real [Osteoarthritis Initiative (OAI)](https://nda.nih.gov/oai) fields — demographics, WOMAC pain/stiffness/function, and per-compartment OARSI radiographic grades (joint space narrowing, osteophytes, sclerosis, attrition) — not the KL label itself

## Test-set results (1,656 held-out X-rays)

| Modality | Accuracy | Quadratic Kappa |
|---|--:|--:|
| Image only (DenseNet121) | 60.3% | 0.792 |
| Clinical Text only (BioClinicalBERT) | 87.4% | 0.953 |
| Multimodal Fusion | 88.9% | 0.959 |

**Caveat:** the clinical-text branch is fed real per-compartment OARSI fields that mechanically
define the KL grade, not free-text patient-reported symptoms — so this is closer to "decode KL's
own defining components written as prose" than "infer severity from symptoms alone." Fusion only
marginally beats text-alone (+1.5pp accuracy) because the text branch already carries near-complete
radiographic signal. Full discussion in the [project README](https://github.com/Darshan-T-P/KneeVision#phase-5--multimodal-fusion-pipeline-).

## Usage

```python
from huggingface_hub import hf_hub_download
import torch

path = hf_hub_download(repo_id="Darshan13/KneeVision-models", filename="best_densenet121.pt")
state = torch.load(path, map_location="cpu", weights_only=False)
```

See [`src/kneevision/models/image_model.py`](https://github.com/Darshan-T-P/KneeVision/blob/main/src/kneevision/models/image_model.py)
(`load_trained_model`) for the full loading logic (backbone reconstruction, ordinal-head detection,
multitask auxiliary-head detection).

## Disclaimer

Research/educational checkpoints only. Not a medical device, not validated for clinical use, not
intended to inform real diagnostic or treatment decisions.