Shahd1sayed/heart-attack-risk-predictor
<h1 align="center">๐ซ Heart Attack Risk Predictor โ Multimodal</h1>
<p align="center"> <strong>An AI-powered clinical decision-support tool that estimates a patient's heart-attack risk from <em>tabular patient data</em>, an <em>ECG image</em>, or <em>both</em> โ using an ensemble-router architecture with a tabular model, an ECG image model, and late-fusion of their scores.</strong> </p>
<p align="center"> <a href="https://huggingface.co/spaces/Shahd1sayed/heart-attack-risk-predictor"><strong>๐ด TRY THE LIVE DEMO ON HUGGING FACE SPACES ๐ด</strong></a> </p>
<p align="center"> <em>โ ๏ธ Educational / portfolio demo โ <strong>not a medical device</strong>. Do not use for real clinical decisions.</em> </p>
๐ Table of Contents
- What's New in v2
- Key Features
- How It Works โ Architecture
- The Two Models
- Tech Stack
- Project Structure
- Getting Started
- Usage & API
- Model Performance
- Testing
- Limitations & Honesty
- Team Members
๐ What's New in v2
Version 1 was a single-model biomarker classifier (8 vitals โ Random Forest). A leakage test showed its ~98% accuracy was largely a biomarker-threshold rule (accuracy fell to ~62% without Troponin & CK-MB). Version 2 re-architects the project into a multimodal ensemble router:
- Two models instead of one โ a tabular model and an ECG image model.
- A router that picks the model(s) based on what the user submits, and averages their scores when both are provided.
- Honest evaluation โ proper metrics (ROC-AUC for the imbalanced tabular task, per-class metrics for the ECG task) and a clear statement of limitations.
The original v1 biomarker research is preserved in research_and_experiments/.โจ Key Features
- Multimodal input โ enter patient data, upload an ECG image, or do both.
- Ensemble router โ one
POST /predictendpoint routes to the right model(s): tabular โ Model A, ECG โ Model B, both โ averaged score. - Missing-data friendly โ blank tabular fields are filled automatically by K-Nearest-Neighbours imputation, so a partial form still works.
- Server-side validation โ out-of-range or non-numeric fields are rejected with a clear message (HTTP 422).
- ECG confidence check โ low-confidence ECG predictions are flagged ("may not be a clear 12-lead ECG").
- Transparent results โ the UI shows each model's score and, in "both" mode, the combined average, so nothing is a black box.
- Modern UI โ dark glassmorphism theme, drag-and-drop ECG upload, color-coded risk badges (๐ด High / ๐ก Moderate / ๐ข Low).
๐ง How It Works โ Architecture
โโโโโโโโโโโโโโโโโ POST /predict (multipart/form-data) โโโโโโโโโโโโโโโโโ
tabular only โโคโ Model A (Framingham: KNN-impute โ scale โ RandomForest) โ p_a โโ โ
โ โโ both โ average โ p โ band (Low/Mod/High)
ECG only โโโโโโคโ Model B (ResNet18 transfer learning on ECG images) โ p_b โโ โ
โ โ
neither โโโโโโโคโ HTTP 400 โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโEach model outputs a scalar risk probability p โ [0, 1]. That value is mapped to a band โ Low (< 0.33), Moderate (0.33โ0.66), High (> 0.66) โ and, when both models run, the two scores are combined by an equal-weight average.
๐ค The Two Models
Model A โ Tabular (Framingham 10-year CHD)
- Data: Framingham Heart Study (4,240 patients, 15 clinical features).
- Target:
TenYearCHDโ probability of coronary heart disease within 10 years. - Pipeline:
StandardScaler โ KNNImputer(k=5) โ classifier, saved as one artifact. - Model selection: RandomForest vs XGBoost by 5-fold CV ROC-AUC โ RandomForest won (0.69 vs 0.67).
- Handles imbalance (~15% positive) with class weighting; the app uses the continuous probability, not a hard 0.5 cutoff.
Model B โ ECG image (ResNet18)
- Data: ECG Images Dataset of Cardiac Patients (928 images, 4 classes).
- Method: transfer learning โ a ResNet18 pretrained on ImageNet, with a new 4-class head; train the head, then fine-tune the last block.
- Classes โ risk weight: Normal
0.0, abnormal heartbeat0.5, post-MI history0.8, myocardial infarction1.0. The 4-class softmax is collapsed to one risk score via these weights.
๐ ๏ธ Tech Stack
๐ Project Structure
heart-attack-risk-predictor/
โ
โโโ app.py # FastAPI app + ensemble router (multipart /predict)
โโโ requirements.txt # Production dependencies
โโโ Dockerfile # Container build (copies app, inference/, models/, static/)
โ
โโโ inference/ # Serving-time prediction package
โ โโโ fusion.py # Risk bands + late-fusion (combine)
โ โโโ framingham.py # Model A inference (with KNN imputation)
โ โโโ ecg.py # Model B inference (softmax โ risk scalar)
โ โโโ validation.py # Server-side field validation
โ
โโโ models/ # Trained artifacts (Git LFS)
โ โโโ framingham_pipeline.joblib # Model A bundle
โ โโโ ecg_resnet.pt # Model B weights
โ โโโ ecg_classes.json # ECG class list + risk weights + preprocessing
โ
โโโ train_framingham.py # Trains Model A
โโโ train_ecg.py # Trains Model B
โ
โโโ static/index.html # Frontend (form + ECG upload + per-branch results)
โ
โโโ DOCUMENTATION.md # Full technical documentation
โโโ DEFENSE_GUIDE.md # Beginner-friendly project defense guide
โ
โโโ research_and_experiments/ # v1 biomarker research (notebook, dataset, old model)๐ Getting Started
Prerequisites
- Python 3.11+ (3.12 recommended)
Install
git clone https://github.com/Shahd1Sayed/heart-attack-risk-predictor.git
cd heart-attack-risk-predictor
pip install -r requirements.txtData (only needed to (re)train โ download from Kaggle)
data/framingham.csv # "Framingham Heart Study dataset"
data/ecg_data/<class>/*.jpg # "ECG Images Dataset of Cardiac Patients"Train (produces the files in models/)
python train_framingham.py
python train_ecg.pyRun
uvicorn app:app --host 127.0.0.1 --port 8000
# open http://127.0.0.1:8000The app boots even before models are trained; a branch needing an untrained model returns HTTP 503 with a hint.
โถ๏ธ Usage & API
POST /predict accepts multipart/form-data with optional tabular fields and an optional ecg image file.
# tabular only
curl -F age=61 -F male=1 -F sysBP=150 -F totChol=240 http://127.0.0.1:8000/predict
# ECG only
curl -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predict
# both (multimodal)
curl -F age=61 -F sysBP=150 -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predictTabular fields: male, age, education, currentSmoker, cigsPerDay, BPMeds, prevalentStroke, prevalentHyp, diabetes, totChol, sysBP, diaBP, BMI, heartRate, glucose (all optional; blanks are KNN-imputed).
Example response (multimodal):
{
"mode": "multimodal",
"risk_level": "High",
"p_risk": 0.7306,
"branches": {
"tabular": { "p_risk": 0.4825, "band": "Moderate",
"detail": {"CHD": 0.4825, "No CHD": 0.5175}, "imputed_fields": ["glucose"] },
"ecg": { "p_risk": 0.9787, "band": "High", "ecg_class": "myocardial_infarction_ecg_images",
"confidence": 0.9587, "low_confidence": false, "warning": null }
}
}Status codes: 200 success ยท 400 no input / unreadable image ยท 422 invalid tabular field ยท 503 model not trained yet.
๐ Model Performance
โ ๏ธ Metric honesty: Framingham is imbalanced (~15% positive), so 85% accuracy is essentially the "always predict no-CHD" baseline โ which is why we lead with ROC-AUC. Predicting a decade ahead from basic clinical features is genuinely hard, so ~0.64โ0.69 AUC is expected. The ECG numbers are strong for a small dataset but optimistic vs. other acquisition setups (the images are photos of printed ECGs).
๐งช Testing
The project ships with a test harness (router paths, response invariants, adversarial inputs, determinism, an ECG serving sweep, input validation, ECG confidence, and concurrency). Result: 62/62 checks pass. See DOCUMENTATION.md for details.
โ๏ธ Limitations & Honesty
- The two models predict different things โ Model A estimates 10-year prognosis; Model B classifies the current ECG. They are trained on different, unpaired populations, so the combined score is a transparent heuristic demonstrating the architecture, not a validated clinical measure.
- The fusion weights are hand-set (equal average) because no paired dataset exists to learn/validate them. A learned fusion on paired data (e.g. PTB-XL) is future work.
- ECG images are photos of printouts โ a small, imbalanced dataset; the CNN may not generalize to other setups.
- No out-of-distribution rejection โ a non-ECG image is still classified (now flagged low-confidence, but not refused).
- Not for clinical use.
