CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes260 downloads2y agoHugging Face02vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3text100K<n<1M0 likes111 downloads1y agoHugging Face03NetherlandsForensicInstitute /wiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M1 likes43 downloads2y agoHugging Face04NetherlandsForensicInstitute /flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use textsentence-similarity100K<n<1M0 likes35 downloads3y agoHugging Face05NetherlandsForensicInstitute /simplewiki-translated-nlThis is a Dutch version of the SimpleWiki text simplification dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes32 downloads2y agoHugging Face06NetherlandsForensicInstitute /stackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes32 downloads2y agoHugging Face07NetherlandsForensicInstitute /msmarco-translated-nlThis is a Dutch version of the MS MARCO dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. A newer translation of this dataset using LLMs is available at NetherlandsForensicInstitute/msmarco-nl. textsentence-similarity100K<n<1M1 likes24 downloads1y agoHugging Face08NetherlandsForensicInstitute /quora-duplicates-translated-nlThis is a Dutch version of the Quora Duplicates dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the Quora Terms of Service. textsentence-similarity100K<n<1M0 likes22 downloads2y agoHugging Face09NetherlandsForensicInstitute /altlex-translated-nlThis is a Dutch version of the AltLex dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M2 likes19 downloads2y agoHugging Face10NetherlandsForensicInstitute /sentence-compression-translated-nlThis is a Dutch version of the Sentence Compression dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes17 downloads2y agoHugging Face11NetherlandsForensicInstitute /coco-captions-translated-nlThis is a Dutch version of the Coco captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.