datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mala-monolingual-filter
MaLA Corpus: Massive Language Adaptation Corpus
This is a cleaned version with some necessary data cleaning.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.mala-monolingual-integration
MaLA Corpus: Massive Language Adaptation Corpus
This is the noisy version that integrates texts from different sources.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.vukuzenzele-monolingual
The Vuk'uzenzele South African Multilingual Corpus
Give Feedback 📑: DSFSI Resource Feedback Form
About Dataset
The dataset was obtained from the South African government magazine Vuk'uzenzele, created by the Government Communication and Information System (GCIS).
The original raw PDFs were obtatined from the Vuk'uzenzele website.
The datasets contain government magazine editions in 11 languages, namely:
Language
Code
Language
Code
English
(eng)
Sepedi
(nso)… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/vukuzenzele-monolingual.mala-monolingual-dedup
MaLA Corpus: Massive Language Adaptation Corpus
This is a deduplicated version after minhash and exact hash deduplication.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.monolingual-tokenizer-dataTodo:
add language to metadata
cite source and explain sampling
monolingual-quechua-iic
Dataset Card for Monolingual-Quechua-IIC
Dataset Summary
We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Supported Tasks and Leaderboards
More Information Needed
Languages
Southern… See the full description on the dataset page: https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic.tibetan_monolingual_A_merged_123_lineszomi-monolingual-corpus
Zomi Monolingual Corpus v1.0
The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated,
reviewed, and permission-approved Zomi sentences. Zomi is represented with the
ISO 639-3 language code ctd (Tedim Chin).
Quick start
from datasets import load_dataset
dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train")
print(dataset.num_rows) # 363401
print(dataset[0]["zomi_text"])
Data fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.swim-ir-monolingual
Dataset Card for SWIM-IR (Monolingual)
This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language.
A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0.
For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website.
What is SWIM-IR?
SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.tibetan_monolingual_Amonolingual_machine_translation_datatibetan_monolingual_A_merged_135_linestibetan_monolingual_A_filteredParaDocs-Monolingualsouth-african-monolingual-corpora-jsonl
South African Languages Pretraining Dataset
This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections.
The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity
Languages Included
Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.monolingual-tokenizer-dataTodo:
add language to metadata
cite source and explain sampling
sanskrit-monolingual-pretraining-corrupted
Dataset Card for "sanskrit-monolingual-pretraining-corrupted"
More Information needed
african-bible-monolingualtibetan_monolingual_gold_merged_135_linesigbo_monolingualA dataset is a collection of Monolingual Igbo sentences.tibetan_monolingual_goldCOVID-19_EU_presscorner_monolingual
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19444
Description
This collection of monolingual documents was generated from content available at https://ec.europa.eu/commission/presscorner.
It includes 1354 documents in total in the following languages: EN 335 DE 266 ES 115 FR 123 IT 273 EL 120 SV 122
Citation
EU presscorner monolingual collections of COVID-19 related documents. (2020, June 15). Version 1.0. [Dataset (Text corpus)].… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_EU_presscorner_monolingual.tibetan_monolingual_S_cleaned_train_test_tokenizedbashkir-wikipedia-monolingual
Bashkir Wikipedia Monolingual Corpus
Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer
training and linguistic research.
Overview
Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801),
cleaned and filtered with automated language identification. The cleaned
configuration is the recommended default for language modelling, tokenization and
linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.kinyarwanda_monolingual_v01.0
!!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!!
!!! This version contains several duplicates and few non-kinyarwanda documents
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.nepali-monolingual-datasettelugu-monolingual-datasettibetan_monolingual_Stibetan_monolingual_S_clean
