stighellemans/meddeid-dutch-synthetic-corpus
MedDeID Dutch synthetic corpus This repository contains 6,493 synthetic Dutch clinical documents with character-offset de-identification spans. It contains no real patient notes or personal information. The corpus was used to train meddeid-dutch-synth. Split policy All 6,493 records are exposed together through one conventional Hugging Face train split. There is no publisher-defined validation split. Here train means the complete model-development corpus; users… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-corpus.
docs: cite MedDeID preprint
Publish canonical patient and caregivers metadata contract
Refresh checksum after dataset card update
Add complete author list and ORCIDs
Stage verified MedDeID publication draft
initial commit
