datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hatecheck-dutch
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-dutch.dutch-colaDutch CoLA is a corpus of linguistic acceptability for Dutch: a dataset consisting of sentences in Dutch, each marked as either acceptable (class 1) or unacceptable (class 0). These sentences are collected from existing descriptions of Dutch grammar (see sources below) with expert-annotated acceptability labels.
Dutch CoLA is part of the group project by students of BA Information Science program at the University of Groningen. List of people involved (alphabetic order):
Abdi, Silvana
Brouwer… See the full description on the dataset page: https://huggingface.co/datasets/GroNLP/dutch-cola.Dutch-GOV-Law-wetten.overheid.nl
Dutch GOV Laws
This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS.
It possible some URLS are missing, around 1% gave a 404 or 405 error.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a short research project for myself.
chatgpt-dutch-simplification
Dataset Card for ChatGPT Dutch Simplification
Dataset Summary
Created in light of a master thesis by Charlotte Van de Velde as part of the Master of Science in Artificial Intelligence at KU Leuven.
Charlotte is supervised by Vincent Vandeghinste and Bram Vanroy.
The dataset contains Dutch source sentences and aligned simplified sentences, generated with ChatGPT. All splits combined, the dataset
consists of 1267 entries.
Charlotte used gpt-3.5-turbo with the following… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification.Dutch-QA-Pairs-RijksoverheidThis dataset originates from the Open Government Web Portal provided by the Dutch Government. It contains Dutch question-answer pairs (vraag-antwoord combinaties or VAC), offering insights into various governmental inquiries and corresponding responses.
Columns: [instruction, input, output, text]
Nr. of entries: 1940
Dutch-Government-Data-for-Bias-detectionDutchCrypticCrosswordThis is a collection of 101 (question,answer) pairs of Dutch cryptic crosswords.
The dataset is derived from copyrighted material and is being used under the EU Copyright Directive's Text and Datamining exception for scientific research.
The dataset is for non-commercial, scientific, and educational use only.
Copyright to NRC / Scryptogram / J.J. Steenhuis.
Zie NRC Handelsblad van 24/8, 31/8, 7/9, 14/9 en 21/9 2024.
For more information see:… See the full description on the dataset page: https://huggingface.co/datasets/MichielBontenbal/DutchCrypticCrossword.Dutch-Speech-Dataset
🎧 Dutch Speech Dataset
The Dutch Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for modern AI and machine learning applications. It includes 179 hours of audio data across 548 files, delivered in MP3 and WAV formats, with a total size of 190 MB. This well-organized audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Dutch-Speech-Dataset.dutch_nursing_home_notes
dutch_nursing_home_notes
Description
This dataset was previously named dutch_nursing_home_records. The name was changed for consistency, as the term ‘notes’ is more commonly used in the community of clinical NLP
This dataset contains synthetic healthcare data generated for NLP experiments.
It mimics real-world client notes of nursing care homes for machine learning and data analysis.
Data Generation
The script describing the data generation can be found here:… See the full description on the dataset page: https://huggingface.co/datasets/ekrombouts/dutch_nursing_home_notes.dutchslang
Dutchslang - 1.0
DutchSlang-dataset is a dataset commited to translating slang (sms/straattaal) to "formal" dutch.
It contatins a few entries with categories, such as region, category, etc..
Warning: There are profanities inside of the dataset, these are listed as "profanity" in the category.
What can this be used for?
Not for training language models directly, but it can be used for:
Autocorrect
Translators
Finetuning Transformers / Tiny transformers that speak… See the full description on the dataset page: https://huggingface.co/datasets/im-lemon/dutchslang.dutch-municipal-sentence-simplification
