CoolFace
Apppublic

Shahd1sayed/heart-attack-risk-predictor

sourceHugging Faceupdated 3mo agoView on Hugging Face
1likes
App README

<h1 align="center">๐Ÿซ€ Heart Attack Risk Predictor โ€” Multimodal</h1>

<p align="center"> <strong>An AI-powered clinical decision-support tool that estimates a patient's heart-attack risk from <em>tabular patient data</em>, an <em>ECG image</em>, or <em>both</em> โ€” using an ensemble-router architecture with a tabular model, an ECG image model, and late-fusion of their scores.</strong> </p>

<p align="center"> <a href="https://huggingface.co/spaces/Shahd1sayed/heart-attack-risk-predictor"><strong>๐Ÿ”ด TRY THE LIVE DEMO ON HUGGING FACE SPACES ๐Ÿ”ด</strong></a> </p>

<p align="center"> <em>โš ๏ธ Educational / portfolio demo โ€” <strong>not a medical device</strong>. Do not use for real clinical decisions.</em> </p>


๐Ÿ“‘ Table of Contents


๐Ÿ†• What's New in v2

Version 1 was a single-model biomarker classifier (8 vitals โ†’ Random Forest). A leakage test showed its ~98% accuracy was largely a biomarker-threshold rule (accuracy fell to ~62% without Troponin & CK-MB). Version 2 re-architects the project into a multimodal ensemble router:

  • โ€”Two models instead of one โ€” a tabular model and an ECG image model.
  • โ€”A router that picks the model(s) based on what the user submits, and averages their scores when both are provided.
  • โ€”Honest evaluation โ€” proper metrics (ROC-AUC for the imbalanced tabular task, per-class metrics for the ECG task) and a clear statement of limitations.
The original v1 biomarker research is preserved in research_and_experiments/.

โœจ Key Features

  • โ€”Multimodal input โ€” enter patient data, upload an ECG image, or do both.
  • โ€”Ensemble router โ€” one POST /predict endpoint routes to the right model(s): tabular โ†’ Model A, ECG โ†’ Model B, both โ†’ averaged score.
  • โ€”Missing-data friendly โ€” blank tabular fields are filled automatically by K-Nearest-Neighbours imputation, so a partial form still works.
  • โ€”Server-side validation โ€” out-of-range or non-numeric fields are rejected with a clear message (HTTP 422).
  • โ€”ECG confidence check โ€” low-confidence ECG predictions are flagged ("may not be a clear 12-lead ECG").
  • โ€”Transparent results โ€” the UI shows each model's score and, in "both" mode, the combined average, so nothing is a black box.
  • โ€”Modern UI โ€” dark glassmorphism theme, drag-and-drop ECG upload, color-coded risk badges (๐Ÿ”ด High / ๐ŸŸก Moderate / ๐ŸŸข Low).

๐Ÿง  How It Works โ€” Architecture

                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ POST /predict (multipart/form-data) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   tabular only โ”€โ”คโ†’ Model A (Framingham: KNN-impute โ†’ scale โ†’ RandomForest) โ†’ p_a โ”€โ”    โ”‚
                 โ”‚                                                                  โ”œโ”€ both โ†’ average โ†’ p โ†’ band (Low/Mod/High)
   ECG only โ”€โ”€โ”€โ”€โ”€โ”คโ†’ Model B (ResNet18 transfer learning on ECG images)      โ†’ p_b โ”€โ”˜    โ”‚
                 โ”‚                                                                       โ”‚
   neither โ”€โ”€โ”€โ”€โ”€โ”€โ”คโ†’ HTTP 400                                                             โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Each model outputs a scalar risk probability p โˆˆ [0, 1]. That value is mapped to a band โ€” Low (< 0.33), Moderate (0.33โ€“0.66), High (> 0.66) โ€” and, when both models run, the two scores are combined by an equal-weight average.


๐Ÿค– The Two Models

Model A โ€” Tabular (Framingham 10-year CHD)

  • โ€”Data: Framingham Heart Study (4,240 patients, 15 clinical features).
  • โ€”Target: TenYearCHD โ€” probability of coronary heart disease within 10 years.
  • โ€”Pipeline: StandardScaler โ†’ KNNImputer(k=5) โ†’ classifier, saved as one artifact.
  • โ€”Model selection: RandomForest vs XGBoost by 5-fold CV ROC-AUC โ€” RandomForest won (0.69 vs 0.67).
  • โ€”Handles imbalance (~15% positive) with class weighting; the app uses the continuous probability, not a hard 0.5 cutoff.

Model B โ€” ECG image (ResNet18)

  • โ€”Data: ECG Images Dataset of Cardiac Patients (928 images, 4 classes).
  • โ€”Method: transfer learning โ€” a ResNet18 pretrained on ImageNet, with a new 4-class head; train the head, then fine-tune the last block.
  • โ€”Classes โ†’ risk weight: Normal 0.0, abnormal heartbeat 0.5, post-MI history 0.8, myocardial infarction 1.0. The 4-class softmax is collapsed to one risk score via these weights.

๐Ÿ› ๏ธ Tech Stack

LayerTechnologyPurpose
BackendFastAPI + UvicornAsync web framework + ASGI server; multipart /predict router
Tabular MLscikit-learn (RandomForest, KNNImputer, StandardScaler), XGBoostModel A pipeline + model comparison
Image MLPyTorch + torchvision (ResNet18)Model B transfer learning
Images / UploadsPillow, python-multipartECG image decoding + file uploads
Datapandas, NumPyData handling
FrontendHTML5, CSS3, Vanilla JSSingle-page glassmorphism UI with ECG drag-and-drop
DeploymentDocker โ†’ Hugging Face SpacesContainerized serving on port 7860

๐Ÿ“ Project Structure

heart-attack-risk-predictor/
โ”‚
โ”œโ”€โ”€ app.py                       # FastAPI app + ensemble router (multipart /predict)
โ”œโ”€โ”€ requirements.txt             # Production dependencies
โ”œโ”€โ”€ Dockerfile                   # Container build (copies app, inference/, models/, static/)
โ”‚
โ”œโ”€โ”€ inference/                   # Serving-time prediction package
โ”‚   โ”œโ”€โ”€ fusion.py                # Risk bands + late-fusion (combine)
โ”‚   โ”œโ”€โ”€ framingham.py            # Model A inference (with KNN imputation)
โ”‚   โ”œโ”€โ”€ ecg.py                   # Model B inference (softmax โ†’ risk scalar)
โ”‚   โ””โ”€โ”€ validation.py            # Server-side field validation
โ”‚
โ”œโ”€โ”€ models/                      # Trained artifacts (Git LFS)
โ”‚   โ”œโ”€โ”€ framingham_pipeline.joblib   # Model A bundle
โ”‚   โ”œโ”€โ”€ ecg_resnet.pt                # Model B weights
โ”‚   โ””โ”€โ”€ ecg_classes.json             # ECG class list + risk weights + preprocessing
โ”‚
โ”œโ”€โ”€ train_framingham.py          # Trains Model A
โ”œโ”€โ”€ train_ecg.py                 # Trains Model B
โ”‚
โ”œโ”€โ”€ static/index.html            # Frontend (form + ECG upload + per-branch results)
โ”‚
โ”œโ”€โ”€ DOCUMENTATION.md             # Full technical documentation
โ”œโ”€โ”€ DEFENSE_GUIDE.md             # Beginner-friendly project defense guide
โ”‚
โ””โ”€โ”€ research_and_experiments/    # v1 biomarker research (notebook, dataset, old model)

๐Ÿš€ Getting Started

Prerequisites

  • โ€”Python 3.11+ (3.12 recommended)

Install

bash
git clone https://github.com/Shahd1Sayed/heart-attack-risk-predictor.git
cd heart-attack-risk-predictor
pip install -r requirements.txt

Data (only needed to (re)train โ€” download from Kaggle)

data/framingham.csv                  # "Framingham Heart Study dataset"
data/ecg_data/<class>/*.jpg          # "ECG Images Dataset of Cardiac Patients"

Train (produces the files in models/)

bash
python train_framingham.py
python train_ecg.py

Run

bash
uvicorn app:app --host 127.0.0.1 --port 8000
# open http://127.0.0.1:8000

The app boots even before models are trained; a branch needing an untrained model returns HTTP 503 with a hint.


โ–ถ๏ธ Usage & API

POST /predict accepts multipart/form-data with optional tabular fields and an optional ecg image file.

bash
# tabular only
curl -F age=61 -F male=1 -F sysBP=150 -F totChol=240 http://127.0.0.1:8000/predict

# ECG only
curl -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predict

# both (multimodal)
curl -F age=61 -F sysBP=150 -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predict

Tabular fields: male, age, education, currentSmoker, cigsPerDay, BPMeds, prevalentStroke, prevalentHyp, diabetes, totChol, sysBP, diaBP, BMI, heartRate, glucose (all optional; blanks are KNN-imputed).

Example response (multimodal):

json
{
  "mode": "multimodal",
  "risk_level": "High",
  "p_risk": 0.7306,
  "branches": {
    "tabular": { "p_risk": 0.4825, "band": "Moderate",
                 "detail": {"CHD": 0.4825, "No CHD": 0.5175}, "imputed_fields": ["glucose"] },
    "ecg": { "p_risk": 0.9787, "band": "High", "ecg_class": "myocardial_infarction_ecg_images",
             "confidence": 0.9587, "low_confidence": false, "warning": null }
  }
}

Status codes: 200 success ยท 400 no input / unreadable image ยท 422 invalid tabular field ยท 503 model not trained yet.


๐Ÿ“Š Model Performance

ModelMetricValue
Model A โ€” FraminghamCV ROC-AUC0.69
Hold-out ROC-AUC0.64
Accuracy0.85 (โ‰ˆ all-negative baseline โ€” see note)
Model B โ€” ECG ResNet18Test accuracy0.90
Macro F10.90
MI precision1.00
โš ๏ธ Metric honesty: Framingham is imbalanced (~15% positive), so 85% accuracy is essentially the "always predict no-CHD" baseline โ€” which is why we lead with ROC-AUC. Predicting a decade ahead from basic clinical features is genuinely hard, so ~0.64โ€“0.69 AUC is expected. The ECG numbers are strong for a small dataset but optimistic vs. other acquisition setups (the images are photos of printed ECGs).

๐Ÿงช Testing

The project ships with a test harness (router paths, response invariants, adversarial inputs, determinism, an ECG serving sweep, input validation, ECG confidence, and concurrency). Result: 62/62 checks pass. See DOCUMENTATION.md for details.


โš–๏ธ Limitations & Honesty

  • โ€”The two models predict different things โ€” Model A estimates 10-year prognosis; Model B classifies the current ECG. They are trained on different, unpaired populations, so the combined score is a transparent heuristic demonstrating the architecture, not a validated clinical measure.
  • โ€”The fusion weights are hand-set (equal average) because no paired dataset exists to learn/validate them. A learned fusion on paired data (e.g. PTB-XL) is future work.
  • โ€”ECG images are photos of printouts โ€” a small, imbalanced dataset; the CNN may not generalize to other setups.
  • โ€”No out-of-distribution rejection โ€” a non-ECG image is still classified (now flagged low-confidence, but not refused).
  • โ€”Not for clinical use.

๐Ÿ‘ฅ Team Members

NameRoleGitHub
Shahd SayedMachine Learning Engineer@Shahd1Sayed
Shahd MohammedFull-Stack Developer