CoolFace
14 results

monolingual

MaLA-LM /mala-monolingual-filter MaLA Corpus: Massive Language Adaptation Corpus This is a cleaned version with some necessary data cleaning. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.text-generation3 likes5.3k downloads2mo agoHugging Facedsfsi /vukuzenzele-monolingual The Vuk'uzenzele South African Multilingual Corpus Give Feedback 📑: DSFSI Resource Feedback Form About Dataset The dataset was obtained from the South African government magazine Vuk'uzenzele, created by the Government Communication and Information System (GCIS). The original raw PDFs were obtatined from the Vuk'uzenzele website. The datasets contain government magazine editions in 11 languages, namely: Language Code Language Code English (eng) Sepedi (nso)… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/vukuzenzele-monolingual.texttranslation1K<n<10K4 likes3.3k downloads3y agoHugging FaceMaLA-LM /mala-monolingual-integration MaLA Corpus: Massive Language Adaptation Corpus This is the noisy version that integrates texts from different sources. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.texttext-generation1B<n<10B2 likes3.1k downloads2mo agoHugging FaceMaLA-LM /mala-monolingual-dedup MaLA Corpus: Massive Language Adaptation Corpus This is a deduplicated version after minhash and exact hash deduplication. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.text-generation2 likes2.9k downloads2mo agoHugging FaceMaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.6k downloads2mo agoHugging Facecatherinearnett /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B1 likes1.4k downloads1y agoHugging Face