CoolFace
Apppublic

Nithins03/clinical-deidentify

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

๐Ÿฅ Clinical-Deidentify: Secure PHI Removal

![CI](https://github.com/sarvanithin/clinical-deidentify/actions)

Fast, regex + transformer hybrid PHI removal for clinical text and documents. Protect patient privacy with clinical-grade accuracy.

UI Mockup

๐Ÿš€ Features

  • โ€”Hybrid Pipeline: Combines deterministic regex for structured PHI (dates, IDs, phones) with state-of-the-art transformers for contextual PHI (patient names, locations).
  • โ€”Expanded Document Support: De-identify PDFs, Word (.docx), and TXT files with a unified interface.
  • โ€”Download Feature: Instantly download de-identified results as a .txt file for safe storage.
  • โ€”Premium Dashboard: A sleek, dark-mode web UI for real-time de-identification and file uploads.
  • โ€”HIPAA Compliant: Docker-native service ensuring all data stays on your infrastructure.
  • โ€”Active Learning: Built-in feedback loop for clinical correction storage.

๐Ÿš€ Quick Start (Docker)

  1. 1.Build:
bash
   docker build -t clinical-deidentify .
  1. 1.Run:
bash
   docker run -d -p 8001:8000 --name clinical-deid-service clinical-deidentify

Dashboard available at: [http://localhost:8001](http://localhost:8001)

Local Installation

  1. 1.Clone & Setup:
bash
   python -m venv venv
   source venv/bin/activate
   pip install -r requirements.txt
  1. 1.Run Server:
bash
   uvicorn app.main:app --reload

Usage

De-identify Single Note

bash
curl -X POST "http://localhost:8000/deidentify" \
     -H "Content-Type: application/json" \
     -d '{"text": "Patient John Doe was admitted on 01/01/2023."}'

Response:

json
{
  "original": "Patient John Doe was admitted on 01/01/2023.",
  "deidentified": "Patient [PATIENT] was admitted on [DATE].",
  "entities": [...]
}

Evaluation

Run the mock benchmarking script:

bash
python eval/evaluate.py

Dataset Benchmarking

The pipeline is designed to be compatible with the 2014 i2b2 de-identification shared task format. You can load i2b2 XML files and map them to the EvalRequest schema within eval/evaluate.py.