monolingual
mala-monolingual-filter
MaLA Corpus: Massive Language Adaptation Corpus
This is a cleaned version with some necessary data cleaning.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.vukuzenzele-monolingual
The Vuk'uzenzele South African Multilingual Corpus
Give Feedback 📑: DSFSI Resource Feedback Form
About Dataset
The dataset was obtained from the South African government magazine Vuk'uzenzele, created by the Government Communication and Information System (GCIS).
The original raw PDFs were obtatined from the Vuk'uzenzele website.
The datasets contain government magazine editions in 11 languages, namely:
Language
Code
Language
Code
English
(eng)
Sepedi
(nso)… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/vukuzenzele-monolingual.mala-monolingual-integration
MaLA Corpus: Massive Language Adaptation Corpus
This is the noisy version that integrates texts from different sources.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.mala-monolingual-dedup
MaLA Corpus: Massive Language Adaptation Corpus
This is a deduplicated version after minhash and exact hash deduplication.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.monolingual-tokenizer-dataTodo:
add language to metadata
cite source and explain sampling
