CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vGassen /Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleantext10K<n<100K0 likes724 downloads1y agoHugging Face02vGassen /Dutch-Basisbestandwetten-Legislation-Lawstext10K<n<100K0 likes290 downloads1y agoHugging Face03stighellemans /meddeid-dutch-synthetic-benchmark MedDeID Dutch synthetic benchmark This repository contains the fixed 300-document synthetic Dutch evaluation benchmark. It contains no real patient notes and must not be mixed into a training or validation partition when reporting MedDeID benchmark results. This is the openly shareable synthetic benchmark described in the manuscript. It is not the separate 300-note hospital benchmark, which contains personal information and is not publicly distributed. Subannotations… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-benchmark.documenttoken-classificationn<1K0 likes212 downloads8d agoHugging Face04schneiderkamplab /dala-dutch-dynaword DaLA Dutch — DynaWord Dutch grammatical acceptability and error correction with synthetic spelling and grammar errors. Provisional, checker-screened training data; not a human-validated gold benchmark. No simplification, paraphrasing or style-transfer task. Configurations 478,916 original/corrupted pairs, 957,832 chat rows per configuration. Every pair contributes a clean control and a corrupted input. The two configurations share sentences and document splits and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dala-dutch-dynaword.texttext-classification1M<n<10M0 likes192 downloads2d agoHugging Face05vGassen /Dutch-Officiele-Publicatiestext100K<n<1M1 likes147 downloads1y agoHugging Face06stighellemans /meddeid-dutch-synthetic-corpus MedDeID Dutch synthetic corpus This repository contains 6,493 synthetic Dutch clinical documents with character-offset de-identification spans. It contains no real patient notes or personal information. The corpus was used to train meddeid-dutch-synth. Split policy All 6,493 records are exposed together through one conventional Hugging Face train split. There is no publisher-defined validation split. Here train means the complete model-development corpus; users… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-corpus.documenttoken-classification1K<n<10K0 likes119 downloads8d agoHugging Face07vGassen /Dutch-Samenwerkende-Catalogitext10K<n<100K0 likes114 downloads1y agoHugging Face08vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3text100K<n<1M0 likes111 downloads1y agoHugging Face09vGassen /Dutch-Staten-Generaal-Digitaal-1814-1995text100K<n<1M0 likes68 downloads1y agoHugging Face10vGassen /Dutch-Tweede-Kamer-APItext10K<n<100K0 likes46 downloads1y agoHugging Face11UMCU /MedQA_DutchTranslation of the English version of MedQA, to Dutch using the GPT 4.1 mini LLM by OpenAI. Attribution If you use this dataset please use the following to credit the creators of MedQA: @article{jin2021disease, title={What disease does this patient have? a large-scale open domain question answering dataset from medical exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={Applied Sciences}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/MedQA_Dutch.tabularquestion-answering10K<n<100K2 likes43 downloads9mo agoHugging Face12UMCU /BioASQ_11B_DutchBio ASQ challenge 11b This can be used to finetune a decoder model for Q/A interaction, alternatively it can be used to create (question, positive, negative) triplets to train a sentence encoder using SBERT. Reference: @inbook{Nentidis_2023, title={Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering}, ISBN={9783031424489}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-031-42448-9_19}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/BioASQ_11B_Dutch.tabularquestion-answering1K<n<10K0 likes39 downloads9mo agoHugging Face13CiviQs /DutchGovBench DutchGovBench v0.1 Evaluation benchmark for Dutch government AI systems. 100 questions across 9 categories, testing knowledge of Dutch law and public administration. What is this? DutchGovBench tests whether AI models can accurately answer questions about Dutch government topics: social support law (Wmo 2015), youth law (Jeugdwet), participation law (Participatiewet), administrative law (Awb), municipal policy, objection procedures, privacy/GDPR, administrative oversight… See the full description on the dataset page: https://huggingface.co/datasets/CiviQs/DutchGovBench.textquestion-answeringn<1K0 likes38 downloads8mo agoHugging Face14UMCU /apollo_english_guidelines_translated_to_dutch_with_gpt4omini Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes33 downloads2y agoHugging Face15careons /dutch-healthcare-pii-ner Dutch Healthcare PII NER Dataset A synthetic Dutch-language Named Entity Recognition (NER) dataset focused on Personally Identifiable Information (PII) detection and anonymization, with a strong emphasis on healthcare contexts in the Netherlands. This dataset is fully synthetic. All texts and entities were generated by AI and do not represent real individuals, organizations, or medical records. Overview Property Value Language Dutch (nl) Samples 400… See the full description on the dataset page: https://huggingface.co/datasets/careons/dutch-healthcare-pii-ner.texttoken-classificationn<1K4 likes31 downloads6mo agoHugging Face16yhavinga /squad_v2_dutch Dataset Card for "squad_v2_dutch" Deprecated: This translation is not recommended. 12% of the translated answers do not appear verbatim in the contexts. Use NetherlandsForensicInstitute/squad-nl-v2.0 instead. Dataset Summary The squad_v2_dutch dataset is a machine-translated version of the SQuAD v2 dataset from English to Dutch. The SQuAD v2 dataset combines the 100,000 questions in SQuAD1.1 with over 50,000 unanswerable questions written adversarially by… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/squad_v2_dutch.textquestion-answering100K<n<1M4 likes30 downloads2y agoHugging Face17vGassen /Dutch-Lokale-Bekendmakingentext1K<n<10K0 likes29 downloads1y agoHugging Face18UMCU /AGCT_Dutch_MariaNMTDutch translation of AGCT using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). If you use this dataset, please cite: @inproceedings{junczys2018marian, title={Marian: Fast Neural Machine Translation in C++}, author={Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and Neckermann, Tom… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/AGCT_Dutch_MariaNMT.tabular100K<n<1M0 likes28 downloads2y agoHugging Face19jjzha /imdb-dutch-instruct Dataset Card for "imdb-dutch-instruct" Dataset Description The original IMBD dataset was translated to Dutch with yhavinga/ul2-large-en-nl. Then, the dataset is converted to an instruct-style dataset with the following templates: The instruction templates: "Is deze recensie positief of negatief?", "Wat is het sentiment van de recensie?", "Wat voor toon heeft de volgende recensie?", "Met wat voor sentiment zou je deze recensie beoordelen?" The target templates: "De… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/imdb-dutch-instruct.text10K<n<100K0 likes27 downloads3y agoHugging Face20CultriX /aya_dutch_dpo_binarized Dataset Card for aya_dutch_dpo This dataset has been created with distilabel. This dataset was created as part of the Data is Better Together project, in particular as part of an ongoing effort to help foster the creation of DPO/ORPO datasets for more languages. The dataset was constructed using the following steps: starting with the aya_dataset and filtering for Dutch examples using the Meta-Llama-3-70B-Instruct model to generate new examples for each promptUsing… See the full description on the dataset page: https://huggingface.co/datasets/CultriX/aya_dutch_dpo_binarized.texttext-generation1K<n<10K1 likes26 downloads2y agoHugging Face21UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face22UMCU /MedQA_SymptomDisease_small_DutchA Dutch translation of this huggingface dataset using GPT4.1 mini, with the courtesy of Prognosis. tabularquestion-answering10K<n<100K0 likes25 downloads9mo agoHugging Face23UMCU /apollo_english_books_translated_to_dutch_with_geminiflash15 Data description Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular100K<n<1M0 likes24 downloads2y agoHugging Face24UMCU /epfl_english_guidelines_translated_to_dutch_with_gpt4omini Dataset Card for Epfl English Guidelines Translated To Dutch With Gpt4Omini This dataset was created by the EPFL, and can found in it original form here The source language: English The original data source: Original Data Source Data description Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini Acknowledgement This is part of the DT4H project with attribution [Cite the paper]. Doi and… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_english_guidelines_translated_to_dutch_with_gpt4omini.tabular10K<n<100K0 likes24 downloads2y agoHugging Face25UMCU /pmcpatient_v2_transformed_to_dutch_discharge_letters_w_geminiflash15gated Dataset Card for Pmcpatient V2 Transformed To Dutch Discharge Letters W Geminiflash15 This dataset was created by: UMCU & Zhengyun Zhao The source language: English The original data source: https://github.com/pmc-patients/pmc-patients Data description We transformed the PMC Patients V2 dataset English discharge letters and translated them to Dutch using Gemini Flash 1.5 Note: all entries with k=0, are English, and all entries with k=1, are Dutch. Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/pmcpatient_v2_transformed_to_dutch_discharge_letters_w_geminiflash15.texttranslation100K<n<1M0 likes23 downloads2y agoHugging Face26aacudad /86k_DUTCH_conversational 🧠 Dutch Instruction Dataset (Generated with Gemini & OpenAI) This dataset was generated using Gemini and OpenAI's API, and is intended for general-purpose Dutch language model training, instruction tuning, and experimentation. Feel free to use it for your own projects or use-cases.If you do, I’d really appreciate it if you could reference or tag me — thanks! 🙌 🚀 Used in DUTCHGPT This dataset has been used to train DUTCHGPT — a fine-tuned version of Gemma and LLaMA… See the full description on the dataset page: https://huggingface.co/datasets/aacudad/86k_DUTCH_conversational.text10K<n<100K0 likes22 downloads1y agoHugging Face27UMCU /apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes21 downloads2y agoHugging Face28tellarin-ai /ntx_llm_inst_dutch Dataset Card for NTX v1 in the Aya format - Dutch subset This dataset is a format conversion for the Dutch data from the original NTX into the Aya instruction format and it's released here under the CC-BY-SA 4.0 license. Dataset Details For the original NTX dataset, the conversion to the Aya instructions format, or more details, please refer to the full dataset in instruction form (https://huggingface.co/datasets/tellarin-ai/ntx_llm_instructions) or to the paper below.… See the full description on the dataset page: https://huggingface.co/datasets/tellarin-ai/ntx_llm_inst_dutch.texttoken-classificationn<1K0 likes20 downloads3y agoHugging Face29CiviQsEU /DutchGovBench DutchGovBench v0.1 A 100-question evaluation benchmark for testing AI models on Dutch government law and policy, covering social support (Wmo 2015), youth care (Jeugdwet), social assistance (Participatiewet), and administrative law (Awb). Purpose DutchGovBench measures whether language models can accurately answer questions about Dutch social legislation. It tests factual knowledge, correct article references, and the ability to handle cross-domain questions, edge cases… See the full description on the dataset page: https://huggingface.co/datasets/CiviQsEU/DutchGovBench.textquestion-answeringn<1K0 likes20 downloads8mo agoHugging Face30aacudad /5K_DUTCH_LEGAL_SUMMARY ⚖️ Dutch Legal Case Dataset (Summarized with Gemini) This dataset consists of 5,000 Dutch legal cases sourced from rechtspraak.nl.Each case includes: The original legal text A summary generated by Gemini The dataset is designed to support long-context training tasks such as legal reasoning and summarization. 🚀 Used in DUTCHGPT This dataset has been used to train DUTCHGPT — a fine-tuned version of Gemma and LLaMA optimized for Dutch. Explore the model here:👉… See the full description on the dataset page: https://huggingface.co/datasets/aacudad/5K_DUTCH_LEGAL_SUMMARY.text1K<n<10K0 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.