CoolFace
Datasetpublic

stighellemans/meddeid-dutch-synthetic-corpus

MedDeID Dutch synthetic corpus This repository contains 6,493 synthetic Dutch clinical documents with character-offset de-identification spans. It contains no real patient notes or personal information. The corpus was used to train meddeid-dutch-synth. Split policy All 6,493 records are exposed together through one conventional Hugging Face train split. There is no publisher-defined validation split. Here train means the complete model-development corpus; users… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-corpus.

sourceHugging Facecc-by-4.0updated 12d agoView on Hugging Face
0likes116downloads
6 commits on main
c42490012d ago

docs: cite MedDeID preprint

stighellemans
cd85ee51mo ago

Publish canonical patient and caregivers metadata contract

stighellemans
7859be12mo ago

Refresh checksum after dataset card update

stighellemans
339f2cb2mo ago

Add complete author list and ORCIDs

stighellemans
a1d25ef2mo ago

Stage verified MedDeID publication draft

stighellemans
76a3a862mo ago

initial commit

stighellemans