datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dutch-Basisbestandwetten-Legislation-Laws-XML-CleanDutch-Basisbestandwetten-Legislation-Lawsmeddeid-dutch-synthetic-benchmark
MedDeID Dutch synthetic benchmark
This repository contains the fixed 300-document synthetic Dutch evaluation
benchmark. It contains no real patient notes and must not be mixed into a
training or validation partition when reporting MedDeID benchmark results.
This is the openly shareable synthetic benchmark described in the manuscript. It is not
the separate 300-note hospital benchmark, which contains personal information
and is not publicly distributed.
Subannotations… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-benchmark.dala-dutch-dynaword
DaLA Dutch — DynaWord
Dutch grammatical acceptability and error correction with synthetic spelling and
grammar errors. Provisional, checker-screened training data; not a human-validated
gold benchmark. No simplification, paraphrasing or style-transfer task.
Configurations
478,916 original/corrupted pairs, 957,832 chat rows
per configuration. Every pair contributes a clean control and a corrupted input.
The two configurations share sentences and document splits and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dala-dutch-dynaword.Dutch-Officiele-Publicatiesmeddeid-dutch-synthetic-corpus
MedDeID Dutch synthetic corpus
This repository contains 6,493 synthetic Dutch clinical documents with
character-offset de-identification spans. It contains no real patient notes or
personal information. The corpus was used to train meddeid-dutch-synth.
Split policy
All 6,493 records are exposed together through one conventional Hugging Face
train split. There is no publisher-defined validation split. Here train
means the complete model-development corpus; users… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-corpus.Dutch-Samenwerkende-CatalogiDutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3Dutch-Staten-Generaal-Digitaal-1814-1995Dutch-Tweede-Kamer-APIMedQA_DutchTranslation of the English version of MedQA,
to Dutch using the GPT 4.1 mini LLM by OpenAI.
Attribution
If you use this dataset please use the following to credit the creators of MedQA:
@article{jin2021disease,
title={What disease does this patient have? a large-scale open domain question answering dataset from medical exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={Applied Sciences}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/MedQA_Dutch.BioASQ_11B_DutchBio ASQ challenge 11b
This can be used to finetune a decoder model for Q/A interaction,
alternatively it can be used to create (question, positive, negative)
triplets to train a sentence encoder using SBERT.
Reference:
@inbook{Nentidis_2023,
title={Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},
ISBN={9783031424489},
ISSN={1611-3349},
url={http://dx.doi.org/10.1007/978-3-031-42448-9_19}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/BioASQ_11B_Dutch.DutchGovBench
DutchGovBench v0.1
Evaluation benchmark for Dutch government AI systems. 100 questions across 9 categories, testing knowledge of Dutch law and public administration.
What is this?
DutchGovBench tests whether AI models can accurately answer questions about Dutch government topics: social support law (Wmo 2015), youth law (Jeugdwet), participation law (Participatiewet), administrative law (Awb), municipal policy, objection procedures, privacy/GDPR, administrative oversight… See the full description on the dataset page: https://huggingface.co/datasets/CiviQs/DutchGovBench.apollo_english_guidelines_translated_to_dutch_with_gpt4omini
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
dutch-healthcare-pii-ner
Dutch Healthcare PII NER Dataset
A synthetic Dutch-language Named Entity Recognition (NER) dataset focused on Personally Identifiable Information (PII) detection and anonymization, with a strong emphasis on healthcare contexts in the Netherlands.
This dataset is fully synthetic. All texts and entities were generated by AI and do not represent real individuals, organizations, or medical records.
Overview
Property
Value
Language
Dutch (nl)
Samples
400… See the full description on the dataset page: https://huggingface.co/datasets/careons/dutch-healthcare-pii-ner.squad_v2_dutch
Dataset Card for "squad_v2_dutch"
Deprecated: This translation is not recommended. 12% of the translated answers do not appear verbatim in the contexts. Use NetherlandsForensicInstitute/squad-nl-v2.0 instead.
Dataset Summary
The squad_v2_dutch dataset is a machine-translated version of the SQuAD v2 dataset from English to Dutch.
The SQuAD v2 dataset combines the 100,000 questions in SQuAD1.1 with over 50,000 unanswerable questions written adversarially by… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/squad_v2_dutch.Dutch-Lokale-BekendmakingenAGCT_Dutch_MariaNMTDutch translation of AGCT using MariaNMT.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
If you use this dataset, please cite:
@inproceedings{junczys2018marian,
title={Marian: Fast Neural Machine Translation in C++},
author={Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and Neckermann, Tom… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/AGCT_Dutch_MariaNMT.imdb-dutch-instruct
Dataset Card for "imdb-dutch-instruct"
Dataset Description
The original IMBD dataset was translated to Dutch with yhavinga/ul2-large-en-nl.
Then, the dataset is converted to an instruct-style dataset with the following templates:
The instruction templates:
"Is deze recensie positief of negatief?",
"Wat is het sentiment van de recensie?",
"Wat voor toon heeft de volgende recensie?",
"Met wat voor sentiment zou je deze recensie beoordelen?"
The target templates:
"De… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/imdb-dutch-instruct.aya_dutch_dpo_binarized
Dataset Card for aya_dutch_dpo
This dataset has been created with distilabel.
This dataset was created as part of the Data is Better Together project, in particular as part of an ongoing effort to help foster the creation of DPO/ORPO datasets for more languages.
The dataset was constructed using the following steps:
starting with the aya_dataset and filtering for Dutch examples
using the Meta-Llama-3-70B-Instruct model to generate new examples for each promptUsing… See the full description on the dataset page: https://huggingface.co/datasets/CultriX/aya_dutch_dpo_binarized.apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
MedQA_SymptomDisease_small_DutchA Dutch translation of this huggingface dataset using GPT4.1 mini, with the courtesy of Prognosis.
apollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
epfl_english_guidelines_translated_to_dutch_with_gpt4omini
Dataset Card for Epfl English Guidelines Translated To Dutch With Gpt4Omini
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with attribution [Cite the paper].
Doi and… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_english_guidelines_translated_to_dutch_with_gpt4omini.pmcpatient_v2_transformed_to_dutch_discharge_letters_w_geminiflash15
Dataset Card for Pmcpatient V2 Transformed To Dutch Discharge Letters W Geminiflash15
This dataset was created by: UMCU & Zhengyun Zhao
The source language: English
The original data source: https://github.com/pmc-patients/pmc-patients
Data description
We transformed the PMC Patients V2 dataset English discharge letters and translated them to Dutch using Gemini Flash 1.5
Note: all entries with k=0, are English, and all entries with k=1, are Dutch.
Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/pmcpatient_v2_transformed_to_dutch_discharge_letters_w_geminiflash15.86k_DUTCH_conversational
🧠 Dutch Instruction Dataset (Generated with Gemini & OpenAI)
This dataset was generated using Gemini and OpenAI's API, and is intended for general-purpose Dutch language model training, instruction tuning, and experimentation.
Feel free to use it for your own projects or use-cases.If you do, I’d really appreciate it if you could reference or tag me — thanks! 🙌
🚀 Used in DUTCHGPT
This dataset has been used to train DUTCHGPT — a fine-tuned version of Gemma and LLaMA… See the full description on the dataset page: https://huggingface.co/datasets/aacudad/86k_DUTCH_conversational.apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
ntx_llm_inst_dutch
Dataset Card for NTX v1 in the Aya format - Dutch subset
This dataset is a format conversion for the Dutch data from the original NTX into the Aya instruction format and it's released here under the CC-BY-SA 4.0 license.
Dataset Details
For the original NTX dataset, the conversion to the Aya instructions format, or more details, please refer to the full dataset in instruction form (https://huggingface.co/datasets/tellarin-ai/ntx_llm_instructions) or to the paper below.… See the full description on the dataset page: https://huggingface.co/datasets/tellarin-ai/ntx_llm_inst_dutch.DutchGovBench
DutchGovBench v0.1
A 100-question evaluation benchmark for testing AI models on Dutch government law and policy, covering social support (Wmo 2015), youth care (Jeugdwet), social assistance (Participatiewet), and administrative law (Awb).
Purpose
DutchGovBench measures whether language models can accurately answer questions about Dutch social legislation. It tests factual knowledge, correct article references, and the ability to handle cross-domain questions, edge cases… See the full description on the dataset page: https://huggingface.co/datasets/CiviQsEU/DutchGovBench.5K_DUTCH_LEGAL_SUMMARY
⚖️ Dutch Legal Case Dataset (Summarized with Gemini)
This dataset consists of 5,000 Dutch legal cases sourced from rechtspraak.nl.Each case includes:
The original legal text
A summary generated by Gemini
The dataset is designed to support long-context training tasks such as legal reasoning and summarization.
🚀 Used in DUTCHGPT
This dataset has been used to train DUTCHGPT — a fine-tuned version of Gemma and LLaMA optimized for Dutch.
Explore the model here:👉… See the full description on the dataset page: https://huggingface.co/datasets/aacudad/5K_DUTCH_LEGAL_SUMMARY.
