CoolFace
Datasetpublic

turkish-nlp-suite/temiz-mC4

Dataset Card for Temiz mC4 Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. split num instances size num of words train 76.432.893 168GB 21.06B Total 76.432.893 168GB 21.06B This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.

sourceHugging Facecc-by-sa-4.0updated 11mo agoView on Hugging Face
2likes546downloads

turkish-nlp-suite/temiz-mC4 · main · files are served by the source, never re-hosted here