stighellemans/meddeid-english-synthetic-benchmark
MedDeID English synthetic clinical benchmark This is a fixed, human-validated benchmark containing 300 synthetic English clinical documents: 150 en-GB and 150 en-US. It contains 1,717 primary PII spans and 7,358 confirmed core-PII subannotation segments. It contains no real patient notes or personal information. Use the entire test split only for final evaluation: from datasets import load_dataset benchmark = load_dataset( "stighellemans/meddeid-english-synthetic-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-english-synthetic-benchmark.
0186
