vkatg/streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.
Update README.md
Rename threshold_sensitivity.csv to data/threshold_sensitivity.csv
Update README.md
Delete EXPERIMENT_REPORT.md
Upload 3 files
Upload 9 files
Upload crossmodal_train.jsonl
Delete latency_summary.csv
Delete policy_metrics.csv
Delete privacy_utility_curve.png
Update README.md
Update README.md
Rename crossmodal_train.jsonl to data/crossmodal_train.jsonl
Upload crossmodal_train.jsonl
Update README.md
Update README.md
Update README.md
Update README.md
Upload 6 files
Update README.md
initial commit
