CoolFace
Datasetpublic

turkish-nlp-suite/temiz-mC4

Dataset Card for Temiz mC4 Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. split num instances size num of words train 76.432.893 168GB 21.06B Total 76.432.893 168GB 21.06B This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.

sourceHugging Facecc-by-sa-4.0updated 11mo agoView on Hugging Face
2likes546downloads
settings

This repository belongs to turkish-nlp-suite on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametemiz-mC4
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerturkish-nlp-suite
Account settings
turkish-nlp-suite/temiz-mC4 · CoolFace