CoolFace
Datasetpublic

Rijgersberg/common_corpus_nl

Common Corpus v2 NL This is a version of Common Corpus v2 filtered to keep only the rows where language is "Dutch". Common Corpus is a very large open and permissible licensed text dataset created by Pleias. Please be sure to acknowledge the creators of the original dataset when using this filtered version. Filtering Common Corpus is a collection of disparate datasets. Note that filtering the entire collection for rows where the language is "Dutch" is not the same… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_nl.

sourceHugging Faceupdated 2y agoView on Hugging Face
4likes1.4kdownloads

Rijgersberg/common_corpus_nl · main · files are served by the source, never re-hosted here