datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3wiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use
simplewiki-translated-nlThis is a Dutch version of the SimpleWiki text simplification dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
stackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
msmarco-translated-nlThis is a Dutch version of the MS MARCO dataset.
Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model,
specifically the huggingface implementation.
A newer translation of this dataset using LLMs is available at NetherlandsForensicInstitute/msmarco-nl.
quora-duplicates-translated-nlThis is a Dutch version of the Quora Duplicates dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the Quora Terms of Service.
altlex-translated-nlThis is a Dutch version of the AltLex dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
sentence-compression-translated-nlThis is a Dutch version of the Sentence Compression dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
coco-captions-translated-nlThis is a Dutch version of the Coco captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
