CoolFace
Datasetpublic

MaLA-LM/mala-monolingual-dedup

MaLA Corpus: Massive Language Adaptation Corpus This is a deduplicated version after minhash and exact hash deduplication. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
2likes2.8kdownloads
settings

This repository belongs to MaLA-LM on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemala-monolingual-dedup
visibilitypublic
licenceodc-by
gatedno
ownerMaLA-LM
Account settings
MaLA-LM/mala-monolingual-dedup · CoolFace