Add metadata and paper link
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -1,375 +1,382 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
<
|
| 13 |
-
<
|
| 14 |
-
-->
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
-
|
| 22 |
-
|
| 23 |
-
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
- [
|
| 28 |
-
- [
|
| 29 |
-
- [
|
| 30 |
-
- [
|
| 31 |
-
- [
|
| 32 |
-
- [
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
-
|
| 49 |
-
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
-
|
| 55 |
-
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
--
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
```
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
|
| 182 |
-
├──
|
| 183 |
-
│
|
| 184 |
-
|
| 185 |
-
│ ├──
|
| 186 |
-
│ ├──
|
| 187 |
-
│
|
| 188 |
-
│
|
| 189 |
-
|
| 190 |
-
│
|
| 191 |
-
├──
|
| 192 |
-
│ ├──
|
| 193 |
-
│
|
| 194 |
-
│
|
| 195 |
-
│
|
| 196 |
-
│
|
| 197 |
-
│
|
| 198 |
-
|
| 199 |
-
│
|
| 200 |
-
│
|
| 201 |
-
│
|
| 202 |
-
│
|
| 203 |
-
└──
|
| 204 |
-
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
├──
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
#
|
| 219 |
-
|
| 220 |
-
|
| 221 |
-
|
| 222 |
-
|
| 223 |
-
|
| 224 |
-
|
| 225 |
-
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
|
| 249 |
-
|
| 250 |
-
-
|
| 251 |
-
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
|
| 258 |
-
|
| 259 |
-
|
| 260 |
-
|
| 261 |
-
|
| 262 |
-
|
| 263 |
-
|
| 264 |
-
|
| 265 |
-
|
| 266 |
-
|
| 267 |
-
|
| 268 |
-
|
| 269 |
-
|
| 270 |
-
--
|
| 271 |
-
|
| 272 |
-
|
| 273 |
-
|
| 274 |
-
|
| 275 |
-
|
| 276 |
-
|
| 277 |
-
|
| 278 |
-
|
| 279 |
-
|
| 280 |
-
|
| 281 |
-
|
| 282 |
-
|
| 283 |
-
|
| 284 |
-
|
| 285 |
-
|
| 286 |
-
|
| 287 |
-
|
| 288 |
-
|
| 289 |
-
|
| 290 |
-
|
| 291 |
-
|
| 292 |
-
|
| 293 |
-
|
| 294 |
-
--
|
| 295 |
-
|
| 296 |
-
|
| 297 |
-
|
| 298 |
-
|
| 299 |
-
|
| 300 |
-
|
| 301 |
-
|
| 302 |
-
|
| 303 |
-
|
| 304 |
-
|
| 305 |
-
|
| 306 |
-
|
| 307 |
-
|
| 308 |
-
|
| 309 |
-
|
| 310 |
-
|
| 311 |
-
|
| 312 |
-
|
| 313 |
-
|
| 314 |
-
|
| 315 |
-
|
| 316 |
-
|
| 317 |
-
|
| 318 |
-
|
| 319 |
-
|
| 320 |
-
|
| 321 |
-
|
| 322 |
-
|
| 323 |
-
|
| 324 |
-
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
-
|
| 328 |
-
|
| 329 |
-
|
| 330 |
-
|
| 331 |
-
|
| 332 |
-
|
| 333 |
-
|
| 334 |
-
|
| 335 |
-
|
| 336 |
-
|
| 337 |
-
|
| 338 |
-
|
| 339 |
-
|
| 340 |
-
|
| 341 |
-
|
| 342 |
-
|
| 343 |
-
|
| 344 |
-
|
| 345 |
-
|
| 346 |
-
|
| 347 |
-
</
|
| 348 |
-
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
|
| 353 |
-
|
| 354 |
-
|
| 355 |
-
|
| 356 |
-
|
| 357 |
-
|
| 358 |
-
|
| 359 |
-
|
| 360 |
-
|
| 361 |
-
|
| 362 |
-
|
| 363 |
-
|
| 364 |
-
|
| 365 |
-
|
| 366 |
-
|
| 367 |
-
|
| 368 |
-
|
| 369 |
-
|
| 370 |
-
|
| 371 |
-
---
|
| 372 |
-
|
| 373 |
-
## <a id="
|
| 374 |
-
|
| 375 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
pipeline_tag: image-classification
|
| 4 |
+
---
|
| 5 |
+
|
| 6 |
+
# ⚡ ECGLight: ECG Digitization & Classification Dashboard
|
| 7 |
+
|
| 8 |
+
This model is described in the paper [ECGLight: Compute-Light Framework For Paper ECG Digitization and Myocardial Infarction Screening](https://huggingface.co/papers/2607.07683).
|
| 9 |
+
|
| 10 |
+
<!--
|
| 11 |
+
<table align="center" border="0" style="border-collapse: collapse; border: none;">
|
| 12 |
+
<tr style="border: none;">
|
| 13 |
+
<td align="center" valign="middle" style="border: none; padding-right: 40px;">
|
| 14 |
+
<img src="assets/logo.png" alt="ECG Digitization & Classification Logo" width="180px" style="border-radius: 14px; box-shadow: 0 4px 16px rgba(15, 23, 42, 0.08);" />
|
| 15 |
+
</td>
|
| 16 |
+
<td align="center" valign="middle" style="border: none;">
|
| 17 |
+
<img src="assets/scai_lab_logo.svg" alt="SCAI Lab Logo" height="100px" />
|
| 18 |
+
</td>
|
| 19 |
+
</tr>
|
| 20 |
+
</table>
|
| 21 |
+
-->
|
| 22 |
+
|
| 23 |
+
An advanced, interactive Streamlit web workstation designed to convert printed/photographed 12-lead paper ECG reports into high-resolution digitized signals and carry out different types of classification over them. The suite is engineered to run with low computational resource requirements, operating seamlessly on a standard consumer laptop GPU (via CUDA) or running completely on CPU.
|
| 24 |
+
|
| 25 |
+
## <a id="table-of-contents"></a>📌 Table of Contents
|
| 26 |
+
|
| 27 |
+
- [🖥️ Web Dashboard Workstation Overview](#web-dashboard-workstation-overview)
|
| 28 |
+
- [🚀 Installation & Setup](#installation-setup)
|
| 29 |
+
- [🧠 Pre-Trained Classifiers & Tasks](#pre-trained-classifiers-tasks)
|
| 30 |
+
- [🚀 Command Line Usage](#command-line-usage)
|
| 31 |
+
- [📁 Repository Structure & Directory Organization](#repository-structure-directory-organization)
|
| 32 |
+
- [📷 How It Works: Signal Digitization](#how-it-works-signal-digitization)
|
| 33 |
+
- [📈 How It Works: Signal Analysis & Visualization](#how-it-works-signal-analysis-visualization)
|
| 34 |
+
- [⚡ How It Works: Heartbeat Segmentation](#how-it-works-heartbeat-segmentation)
|
| 35 |
+
- [🧠 How It Works: Cardiac Classification](#how-it-works-cardiac-classification)
|
| 36 |
+
- [🤝 Collaborating Institutions](#collaborating-institutions)
|
| 37 |
+
- [📄 Citation](#citation)
|
| 38 |
+
- [👥 Authors & Contact](#authors-contact)
|
| 39 |
+
- [📄 License](#license)
|
| 40 |
+
|
| 41 |
+
## <a id="web-dashboard-workstation-overview"></a>🖥️ Web Dashboard Workstation Overview
|
| 42 |
+
|
| 43 |
+
The dashboard provides a premium, responsive user interface designed for research, education, and clinical workflow exploration. It coordinates the digitization and classification pipelines into a unified, lightweight web application.
|
| 44 |
+
|
| 45 |
+
### Workstation Modules
|
| 46 |
+
|
| 47 |
+
1. **📷 ECG Image Digitizer**:
|
| 48 |
+
- Upload any scanned or photographed ECG image (`.png`, `.jpg`, `.jpeg`).
|
| 49 |
+
- Run the sequential YOLOv11 pipeline step-by-step with real-time progress indicators.
|
| 50 |
+
- Outputs a summary of detected leads and total samples.
|
| 51 |
+
- Automatically saves the digitized CSV to disk under `output/digitization/latest_digitized.csv` for downstream consumption.
|
| 52 |
+
|
| 53 |
+
2. **📈 ECG Signal Viewer**:
|
| 54 |
+
- Visualizes multi-channel ECG signals interactively using client-side native Streamlit line charts (supporting zoom, pan, and hover tooltips).
|
| 55 |
+
- Supports stacked subplots (with distinct clinical colors for each lead: clinical red, teal, deep blue, yellow, purple, etc.) or overlaid graphs.
|
| 56 |
+
- Displays statistical summaries (mean, standard deviation, min/max, range) and allows row-by-row signal previewing.
|
| 57 |
+
|
| 58 |
+
3. **❤️ ECG Classification**:
|
| 59 |
+
- Predicts cardiac conditions using pre-trained ensemble and deep learning classifiers.
|
| 60 |
+
- Automatically segments raw signals into heartbeats around R-peaks using the Pan-Tompkins algorithm before running inference.
|
| 61 |
+
- **Inference Mode (No Ground-Truth)**: If the uploaded CSV lacks diagnostic labels, the dashboard displays a downloadable **Predictions Table** detailing predicted class diagnoses and model confidence probabilities.
|
| 62 |
+
- **Evaluation Mode (With Ground-Truth)**: If labels are present, the page calculates and plots performance metrics (Accuracy, F1-Score, Sensitivity, Specificity, Confusion Matrix).
|
| 63 |
+
|
| 64 |
+
### 🔄 Workflow
|
| 65 |
+
|
| 66 |
+
The workstation coordinates the pipeline through four distinct steps: **Digitization**, **Analysis**, **Segmentation**, and **Classification**. The technical workflows for each process are documented below in their respective **How It Works** sections.
|
| 67 |
+
|
| 68 |
+
---
|
| 69 |
+
|
| 70 |
+
## <a id="installation-setup"></a>🚀 Installation & Setup
|
| 71 |
+
|
| 72 |
+
### Prerequisites
|
| 73 |
+
- Python 3.9
|
| 74 |
+
- CUDA-capable GPU recommended (automatically falls back to CPU if unavailable).
|
| 75 |
+
|
| 76 |
+
### Conda Environment & Model Setup
|
| 77 |
+
|
| 78 |
+
1. Clone the repository and navigate to the project directory:
|
| 79 |
+
```bash
|
| 80 |
+
git clone https://github.com/scai-lab/ECG-Digitization-Classification.git
|
| 81 |
+
cd ECG-Digitization-Classification
|
| 82 |
+
```
|
| 83 |
+
|
| 84 |
+
2. Create the conda environment using the provided `environment.yml` configuration:
|
| 85 |
+
```bash
|
| 86 |
+
conda env create -f environment.yml
|
| 87 |
+
conda activate infer
|
| 88 |
+
```
|
| 89 |
+
|
| 90 |
+
> [!IMPORTANT]
|
| 91 |
+
> **Windows Compatibility & TensorFlow Setup**:
|
| 92 |
+
> If you are on Windows and encounter native runtime loading failures (`ImportError: DLL load failed while importing _pywrap_tensorflow_internal: A dynamic link library (DLL) initialization routine failed`), you need to install a stable version pairing of TensorFlow and Protobuf:
|
| 93 |
+
> ```bash
|
| 94 |
+
> pip install tensorflow==2.15.0 protobuf==4.25.3
|
| 95 |
+
> ```
|
| 96 |
+
> *(Make sure no background Streamlit or Python tasks are running when executing this command, to prevent file locking issues on `.pyd` libraries).*
|
| 97 |
+
|
| 98 |
+
3. **Download Pre-Trained Model Weights**:
|
| 99 |
+
Due to their file sizes, the YOLO detection checkpoints and pre-trained classifiers are hosted externally. Download the `models/` directory from the link below and place it directly in the root of the project:
|
| 100 |
+
|
| 101 |
+
👉 **[Download Pre-Trained Models Directory (ETH Zürich Polybox)](https://polybox.ethz.ch/index.php/s/GDACstPtsoTrrWH)**
|
| 102 |
+
|
| 103 |
+
Once extracted, verify that the weights are located inside the directory tree structure:
|
| 104 |
+
```text
|
| 105 |
+
models/
|
| 106 |
+
├── digitization_models/
|
| 107 |
+
│ ├── yolo11_full/weights/best.pt
|
| 108 |
+
│ ├── yolo11_lead/weights/best.pt
|
| 109 |
+
│ ├── yolo11_pulse/weights/best.pt
|
| 110 |
+
│ └── yolo11_patch/weights/best.pt
|
| 111 |
+
└── classifier_models/
|
| 112 |
+
├── mi_vs_normal_segmented/
|
| 113 |
+
├── omi_vs_nonomi/
|
| 114 |
+
└── ecg_surgery/
|
| 115 |
+
```
|
| 116 |
+
|
| 117 |
+
Key packages installed by the environment: `torch 2.7`, `ultralytics 8.3`, `opencv-python 4.11`, `scikit-image 0.24`, `wfdb 4.3`, `patched-yolo-infer 1.3.8`, `sktime`, `streamlit`.
|
| 118 |
+
|
| 119 |
+
---
|
| 120 |
+
|
| 121 |
+
## <a id="pre-trained-classifiers-tasks"></a>🧠 Pre-Trained Classifiers & Tasks
|
| 122 |
+
|
| 123 |
+
The classification engine supports three diagnostic tasks using the pre-trained weights in `classifier_models/`:
|
| 124 |
+
|
| 125 |
+
| Classification Task | Model Type | Expected Input Shape | Test Accuracy | Positive Class |
|
| 126 |
+
| :--- | :--- | :--- | :---: | :--- |
|
| 127 |
+
| **Normal vs Myocardial Infarction (MI) - Segmented** | Arsenal | 12 leads × 140 timesteps | **92.3%** | `MYOCARDIAL_INFARCTION` |
|
| 128 |
+
| **Occlusive MI (OMI) vs non-OMI** | Rocket | 12 leads × 141 timesteps | **88.9%** | `OMI` |
|
| 129 |
+
| **Pre-Procedural vs Post-Procedural MI** | InceptionTime | 12 leads × 140 timesteps | **91.4%** | `pre-procedural MI` |
|
| 130 |
+
|
| 131 |
+
- **Arsenal**: An ensemble of ROCKET classifiers utilizing random convolutional kernels to extract feature representations combined with ridge regression.
|
| 132 |
+
- **Rocket**: Random Omni-directional Kernel Extraction (ROCKET) classifier, computing kernel convolutions quickly for high-dimensional time-series data.
|
| 133 |
+
- **InceptionTime**: A deep convolutional network ensemble modeled on the Inception architecture, extracting multi-scale temporal features.
|
| 134 |
+
|
| 135 |
+
---
|
| 136 |
+
|
| 137 |
+
## <a id="command-line-usage"></a>🚀 Command Line Usage
|
| 138 |
+
|
| 139 |
+
### Run Batch Digitization (`run_org.py`)
|
| 140 |
+
|
| 141 |
+
The batch processing script processes nested hospital directories, exporting structured folders of digitized CSVs:
|
| 142 |
+
|
| 143 |
+
1. Configure path variables at the top of `run_org.py`:
|
| 144 |
+
```python
|
| 145 |
+
ORGANIZED_DIR = "../ecg_files/ECG_organized_all" # Input dataset root
|
| 146 |
+
OUTPUT_DIR = "../ecg_files/ECG_digitized" # Mirrored CSV directory
|
| 147 |
+
CATEGORIES = ["pre", "index", "post"] # Categories to process
|
| 148 |
+
```
|
| 149 |
+
|
| 150 |
+
2. Run the script:
|
| 151 |
+
```bash
|
| 152 |
+
python run_org.py
|
| 153 |
+
```
|
| 154 |
+
|
| 155 |
+
### Run Model Inference (`run_inference.py`)
|
| 156 |
+
|
| 157 |
+
Execute predictions directly on digitized data from the command line using `run_inference.py` located in `archive/classification/`:
|
| 158 |
+
|
| 159 |
+
```bash
|
| 160 |
+
# MI vs Normal Segmented heartbeat classification
|
| 161 |
+
python archive/classification/run_inference.py --model mi_vs_normal_segmented --input data/ptb_xl/segmented_heartbeats.csv
|
| 162 |
+
|
| 163 |
+
# OMI vs non-OMI classification
|
| 164 |
+
python archive/classification/run_inference.py --model omi_vs_nonomi --input data/ecg_matrix_omi_segmented_50_150_90.csv
|
| 165 |
+
|
| 166 |
+
# Custom output file path
|
| 167 |
+
python archive/classification/run_inference.py --model ecg_surgery --input data/ecg_surgery_segmented_50_150_70.csv --output results/surgery_preds.csv
|
| 168 |
+
```
|
| 169 |
+
|
| 170 |
+
---
|
| 171 |
+
|
| 172 |
+
## <a id="repository-structure-directory-organization"></a>📁 Repository Structure & Directory Organization
|
| 173 |
+
|
| 174 |
+
The repository is structured to maintain a clean root directory, moving utility runners, UI views, model checkpoints, and legacy/training scripts into distinct modules:
|
| 175 |
+
|
| 176 |
+
```
|
| 177 |
+
.
|
| 178 |
+
├── app.py # Streamlit application main router
|
| 179 |
+
├── config.py # Centralized configuration and model registry
|
| 180 |
+
├── digitization.py # Core ECGImage extraction pipeline class
|
| 181 |
+
├── environment.yml # Conda environment dependency file
|
| 182 |
+
├── README.md # Comprehensive repository documentation
|
| 183 |
+
│
|
| 184 |
+
├── backend/ # Dashboard background execution adapters
|
| 185 |
+
│ ├── __init__.py # Backend package declaration
|
| 186 |
+
│ ├── digitization_runner.py # YOLO loader and single-image processor
|
| 187 |
+
│ └── classification_runner.py # Pre-trained model loader and preprocessor
|
| 188 |
+
│
|
| 189 |
+
├── utils/ # Streamlit front-end page components
|
| 190 |
+
│ ├── __init__.py # Utils package declaration
|
| 191 |
+
│ ├── branding.py # Sidebar titles, headers, and footer logos
|
| 192 |
+
│ ├── css.py # Custom clinical theme and grid background CSS
|
| 193 |
+
│ ├── hardware.py # Displays CPU/GPU hardware properties (cached)
|
| 194 |
+
│ ├── page_digitizer.py # Front-end for the ECG Digitizer page
|
| 195 |
+
│ ├── page_csv_viewer.py # Front-end for the interactive Signal Viewer
|
| 196 |
+
│ └── page_classifier.py # Front-end for the Classification workstation
|
| 197 |
+
│
|
| 198 |
+
├── models/ # Relocated YOLO checkpoints and classifiers
|
| 199 |
+
│ ├── digitization_models/ # YOLO v11 checkpoints for digitization
|
| 200 |
+
│ │ ├── yolo11_full/ # YOLO Bounding boxes
|
| 201 |
+
│ │ ├── yolo11_lead/ # YOLO Lead names
|
| 202 |
+
│ │ ├── yolo11_pulse/ # YOLO Reference pulses
|
| 203 |
+
│ │ └── yolo11_patch/ # YOLO Waveform segmentations
|
| 204 |
+
│ │
|
| 205 |
+
│ └── classifier_models/ # Bundled pre-trained diagnostic classifiers
|
| 206 |
+
│ ├── mi_vs_normal_segmented/ # Pre-trained Arsenal model (segmented beats)
|
| 207 |
+
│ ├── omi_vs_nonomi/ # Pre-trained Rocket model (segmented beats)
|
| 208 |
+
│ └── ecg_surgery/ # Pre-trained InceptionTime model (segmented beats)
|
| 209 |
+
│
|
| 210 |
+
└── archive/ # Archived developer, training, and legacy scripts
|
| 211 |
+
└── classification/
|
| 212 |
+
├── train_and_save_models.py# Script used to compile pre-trained models
|
| 213 |
+
├── run_inference.py # Independent CLI inference execution script
|
| 214 |
+
├── run_classification.py # Baseline MLP classifier pipeline
|
| 215 |
+
├── run_benchmarking.py # Comparison benchmarking suite
|
| 216 |
+
├── run_lead_importance_test.py# Individual lead performance evaluator
|
| 217 |
+
├── feature_analysis.py # Original all-in-one analysis script
|
| 218 |
+
├── aggregate_subject_metrics.py# Multi-subject performance aggregator
|
| 219 |
+
├── dataset_curate.py # Local dataset curation utility
|
| 220 |
+
└── re_plotter.py # Advanced Gaussian signal generator & visualizer
|
| 221 |
+
```
|
| 222 |
+
|
| 223 |
+
---
|
| 224 |
+
|
| 225 |
+
## <a id="how-it-works-signal-digitization"></a>📷 How It Works: Signal Digitization
|
| 226 |
+
|
| 227 |
+
The core class [digitization.py](file:///d:/Projects/ECGLight/digitization.py) operates a multi-stage sequential computer vision pipeline to translate raster images into digitized signals:
|
| 228 |
+
|
| 229 |
+
```mermaid
|
| 230 |
+
graph TD
|
| 231 |
+
A[ECG Image Upload] --> B[Preprocessing: Otsu & Blurring]
|
| 232 |
+
B --> C[YOLOv11 Detection & Segmentation]
|
| 233 |
+
subgraph YOLOv11 Models
|
| 234 |
+
C1[yolo11_full: Lead Boundaries]
|
| 235 |
+
C2[yolo11_lead: Text Name Labels]
|
| 236 |
+
C3[yolo11_pulse: Calibration Pulses]
|
| 237 |
+
C4[yolo11_patch: Waveform Segments]
|
| 238 |
+
end
|
| 239 |
+
C --> C1 & C2 & C3 & C4
|
| 240 |
+
C1 & C2 & C3 & C4 --> D[Hough Lines Calibration]
|
| 241 |
+
D --> E[K-Means Row & Column Grid Construction]
|
| 242 |
+
E --> F[Anti-Leakage Connected Components Filter]
|
| 243 |
+
F --> G[Centroid Trace & Resampling to 500Hz]
|
| 244 |
+
G --> H[Export latest_digitized.csv]
|
| 245 |
+
```
|
| 246 |
+
|
| 247 |
+
1. **Preprocessing**: Cleans the scanned image using shadow-removal masks, Otsu binarization, and Gaussian blurring to isolate ink lines from paper textures.
|
| 248 |
+
2. **YOLO Segmentation**: Applies a patched YOLO segmentation model at three crop scales (`4×`, `4.5×`, and `5×` height) to isolate individual lead waveform contours.
|
| 249 |
+
3. **Sequential Detections**: Runs three YOLO models in parallel:
|
| 250 |
+
- `yolo11_full`: Bounding boxes for the 12 lead channels.
|
| 251 |
+
- `yolo11_lead`: Text labels representing lead names (I, II, aVR...).
|
| 252 |
+
- `yolo11_pulse`: Bounding boxes for the calibration reference pulses (typically 1mV high, representing vertical scale).
|
| 253 |
+
4. **Scale Calibration**: Fits Hough lines to the calibration pulse boundaries. The pixel height determines the voltage scale (`volt/pixel`), while the width determines the time scale (`time/pixel`).
|
| 254 |
+
5. **Grid Construction**: Employs K-Means clustering on lead coordinates to map rows and columns, automatically parsing standard Cabrera orders and grid formats (3×4, 4×3, 6×2, 12×1).
|
| 255 |
+
6. **Signal Extraction & Post-Processing**: Traces contours to extract raw pixel centroids, performs baseline correction, applies linear interpolation to bridge gaps, and resamples to a standard **500 Hz** frequency calibrated in **millivolts (mV)**. In addition, an **anti-leakage component filter** is executed per cell crop using connected components analysis to automatically identify the primary waveform trace and strip out smaller, boundary-adjacent components (leaked signals from neighboring leads) that sit far from the row baseline.
|
| 256 |
+
|
| 257 |
+
---
|
| 258 |
+
|
| 259 |
+
## <a id="how-it-works-signal-analysis-visualization"></a>📈 How It Works: Signal Analysis & Visualization
|
| 260 |
+
|
| 261 |
+
Once continuous signals are extracted, the dashboard runs analytical tasks and displays interactive previews:
|
| 262 |
+
|
| 263 |
+
```mermaid
|
| 264 |
+
graph TD
|
| 265 |
+
A[Upload Digitized CSV] --> B[Parse Lead Voltages & Timestamps]
|
| 266 |
+
B --> C[Vega-Lite Interactive Visualizer]
|
| 267 |
+
C --> C1[Render Stacked Leads]
|
| 268 |
+
C --> C2[Render Overlaid Signals]
|
| 269 |
+
B --> D[Compute Signal Statistics: Mean, SD, Min/Max]
|
| 270 |
+
D --> E[Display Summary Dataframes & Row Previews]
|
| 271 |
+
```
|
| 272 |
+
|
| 273 |
+
1. **Parse Signals**: Reads digitized CSV format, validating lead names and timestamps.
|
| 274 |
+
2. **Vega-Lite Visualization**: Renders interactive charts supporting native client-side zoom, pan, and hover tooltips for all channels.
|
| 275 |
+
3. **Signal Statistics**: Automatically computes statistical characteristics (mean, standard deviation, min/max values) for each lead.
|
| 276 |
+
|
| 277 |
+
---
|
| 278 |
+
|
| 279 |
+
## <a id="how-it-works-heartbeat-segmentation"></a>⚡ How It Works: Heartbeat Segmentation
|
| 280 |
+
|
| 281 |
+
To prepare continuous digitized signals for the classification models, the pipeline runs the Pan-Tompkins R-peak detection algorithm:
|
| 282 |
+
|
| 283 |
+
```mermaid
|
| 284 |
+
graph TD
|
| 285 |
+
A[Digitized 500Hz Signal] --> B[Bandpass Filter 5-15Hz]
|
| 286 |
+
B --> C[Derivative Filter]
|
| 287 |
+
C --> D[Squaring Operation]
|
| 288 |
+
D --> E[Moving Window Integration]
|
| 289 |
+
E --> F[Adaptive Thresholding & R-Peak Search]
|
| 290 |
+
F --> G[Extract 140-sample Beats: 50ms pre-R, 150ms post-R]
|
| 291 |
+
G --> H[Max-Absolute Voltage Normalization]
|
| 292 |
+
```
|
| 293 |
+
|
| 294 |
+
1. **Filtering**: The Lead II signal is filtered via a bandpass filter (5-15 Hz) to suppress muscle noise, baseline wander, and T-wave interference.
|
| 295 |
+
2. **Differentiation**: Computes the slope of the signal to highlight the rapid change in the QRS complex.
|
| 296 |
+
3. **Squaring**: Performs point-by-point squaring to amplify QRS slopes while attenuating smaller waves.
|
| 297 |
+
4. **Integration**: A moving window integrator (typically 150ms wide) compiles the slope information into a peak window.
|
| 298 |
+
5. **Adaptive Thresholding & Peak Search**: Dynamically computes threshold constants based on average noise and signal levels, locating R-peaks.
|
| 299 |
+
6. **Beat Windowing**: Extracts a localized heartbeat around each R-peak (typically extending 50ms before and 150ms after the peak), normalizes the voltage per heartbeat using max-absolute scaling, and truncates/pads the resulting segments to the target model input width (e.g. 140 or 141 timesteps).
|
| 300 |
+
|
| 301 |
+
---
|
| 302 |
+
|
| 303 |
+
## <a id="how-it-works-cardiac-classification"></a>🧠 How It Works: Cardiac Classification
|
| 304 |
+
|
| 305 |
+
The heartbeat segment tensors are evaluated using pre-trained time-series classification models:
|
| 306 |
+
|
| 307 |
+
```mermaid
|
| 308 |
+
graph TD
|
| 309 |
+
A[Segmented Heartbeats] --> B[Numpy3D Reshaping: N_instances × 12_leads × N_timesteps]
|
| 310 |
+
B --> C[Select Classification Task]
|
| 311 |
+
subgraph Model Registry
|
| 312 |
+
C1[Normal vs MI: Arsenal]
|
| 313 |
+
C2[OMI vs non-OMI: Rocket]
|
| 314 |
+
C3[Pre vs Post-Procedural MI: InceptionTime]
|
| 315 |
+
end
|
| 316 |
+
C --> C1 & C2 & C3
|
| 317 |
+
C1 & C2 & C3 --> D[Load Pre-Trained Pickled Estimator]
|
| 318 |
+
D --> E[Predict Class Labels & Probabilities]
|
| 319 |
+
E --> F[Generate Downloadable Predictions CSV]
|
| 320 |
+
```
|
| 321 |
+
|
| 322 |
+
1. **Numpy3D Formatting**: Formats the heartbeat segments into a standard `sktime` `Numpy3D` tensor with shape `(N_instances, 12_leads, N_timesteps)`.
|
| 323 |
+
2. **Dynamic Task Selection**: Loads the pre-trained pickled model corresponding to the selected classification task.
|
| 324 |
+
3. **Model Inference**: Evaluates the model to compute class predictions and probability confidences.
|
| 325 |
+
4. **Result Generation**: Automatically builds downloadable prediction tables and calculates performance metrics if ground-truth labels are present in the dataset.
|
| 326 |
+
|
| 327 |
+
---
|
| 328 |
+
|
| 329 |
+
## <a id="collaborating-institutions"></a>🤝 Collaborating Institutions
|
| 330 |
+
|
| 331 |
+
This project was developed in collaboration with:
|
| 332 |
+
|
| 333 |
+
- [ETH Zürich](https://ethz.ch)
|
| 334 |
+
- [Istituto Cardiocentro Ticino (EOC)](https://www.cardiocentro.org)
|
| 335 |
+
- [Università della Svizzera italiana (USI)](https://www.usi.ch)
|
| 336 |
+
- [Università della Campania Luigi Vanvitelli](https://www.unicampania.it)
|
| 337 |
+
|
| 338 |
+
<p align="center">
|
| 339 |
+
<a href="https://ethz.ch" target="_blank">
|
| 340 |
+
<img src="assets/ETH_Zürich_Logo_black.svg.png" alt="ETH Zürich" height="30px" style="vertical-align: middle; margin: 0 15px;" />
|
| 341 |
+
</a>
|
| 342 |
+
|
| 343 |
+
<a href="https://www.cardiocentro.org" target="_blank">
|
| 344 |
+
<img src="assets/eoc_logo.png" alt="Istituto Cardiocentro Ticino (EOC)" height="35px" style="vertical-align: middle; margin: 0 15px;" />
|
| 345 |
+
</a>
|
| 346 |
+
|
| 347 |
+
<a href="https://www.usi.ch" target="_blank">
|
| 348 |
+
<img src="assets/usi_logo.png" alt="USI" height="35px" style="vertical-align: middle; margin: 0 15px;" />
|
| 349 |
+
</a>
|
| 350 |
+
|
| 351 |
+
<a href="https://www.unicampania.it" target="_blank">
|
| 352 |
+
<img src="assets/Logo_Vanvitelli_university.svg.png" alt="Università della Campania Luigi Vanvitelli" height="35px" style="vertical-align: middle; margin: 0 15px;" />
|
| 353 |
+
</a>
|
| 354 |
+
</p>
|
| 355 |
+
|
| 356 |
+
---
|
| 357 |
+
|
| 358 |
+
## <a id="citation"></a>📄 Citation
|
| 359 |
+
|
| 360 |
+
```bibtex
|
| 361 |
+
@article{natraj2026ecglight,
|
| 362 |
+
title={ECGLight: Compute-Light Framework For Paper ECG Digitization and Myocardial Infarction Screening},
|
| 363 |
+
author={Natraj, Shreyasvi and Achtari, Cyrus and Gragnano, Felice and Milzi, Andrea and Valgimigli, Marco and Paez-Granados, Diego},
|
| 364 |
+
journal={arXiv preprint arXiv:2607.07683},
|
| 365 |
+
year={2026},
|
| 366 |
+
url={https://arxiv.org/abs/2607.07683},
|
| 367 |
+
doi={10.48550/arXiv.2607.07683}
|
| 368 |
+
}
|
| 369 |
+
```
|
| 370 |
+
|
| 371 |
+
---
|
| 372 |
+
|
| 373 |
+
## <a id="authors-contact"></a>👥 Authors & Contact
|
| 374 |
+
|
| 375 |
+
- **Shreyasvi Natraj** — [snatraj@ethz.ch](mailto:snatraj@ethz.ch)
|
| 376 |
+
- **Cyrus Achtari**
|
| 377 |
+
|
| 378 |
+
---
|
| 379 |
+
|
| 380 |
+
## <a id="license"></a>📄 License
|
| 381 |
+
|
| 382 |
+
This project is released under the **Non-Commercial Academic and Research License Agreement**. Please refer to the [LICENSE](file:///d:/Projects/ECGLight/LICENSE) file in the repository root for the full licensing terms. The codebase and trained model weights are provided free of charge for personal, academic, and non-profit research use only. Commercial use is strictly prohibited.
|