turkish-nlp-suite/temiz-mC4
Dataset Card for Temiz mC4 Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. split num instances size num of words train 76.432.893 168GB 21.06B Total 76.432.893 168GB 21.06B This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face