datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
low_resource_languages_pretrain_data5low_resource_languages_pretrain_data2low_resource_languages_pretrain_data8low_resource_languages_pretrainLow-resource-QE-DA-dataset
Low-resource QE-DA Dataset
Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE.
Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv)
Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.low_resource_languages_pretrain_data4low_resource_languages_pretrain_datalow-resource-rag-indexes
Low-Resource RAG: Wikipedia FAISS Indexes (BGE-M3)
FAISS indexes over Wikipedia 2023 passages for four languages, embedded with BAAI/bge-m3. Source corpus: CohereLabs/wikipedia-2023-11-embed-multilingual-v3.
Each language folder has two files aligned row-by-row:
index.faiss — FAISS IndexFlatIP, dim=1024, L2-normalized BGE-M3 vectors
dataset/ — Arrow dataset with _id, title, text per passage
Languages
Language
Code
Passages
Size
Hindi
hi
TBD
2.6G… See the full description on the dataset page: https://huggingface.co/datasets/Pika4028/low-resource-rag-indexes.Urdu-Low-Resource-Language-Dataset
Dataset Card for Low-Resource-Language-Dataset
This dataset is designed to aid Natural Language Processing (NLP) research on low-resource languages, particularly Urdu. It includes structured datasets and preprocessing tools curated from the BBC Urdu website.
Dataset Details
Dataset Description
This dataset contains articles, summaries, and topics scraped from BBC Urdu. It is structured into training and testing datasets for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Subayyal/Urdu-Low-Resource-Language-Dataset.rumantsch-varieties-sentencesmgsm-all-low-resource-translatedAdaption-low-resource-doc-qa
Adaption Low-Resource Document Q/A
This dataset is a remastered version of
Reubencf/magazines-multilingual-vqa
prepared using Adaption's Adaptive Data platform,
with a deliberate focus on low-resource source languages — the
languages that are underrepresented in most open multimodal datasets.
What's inside
10,200 rows of multilingual document question-answer pairs grounded in
public-domain magazine / newspaper pages from archive.org.
Every row carries verbatim OCR in… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-doc-qa.kenyan-low-resource-language-data
Low-Resource Language Data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin, and Dholuo
Description
This dataset consists of three parallel corpora:
Kidaw'ida (Dawida)-Kiswahili (dav_swa)
Kalenjin-Kiswahili (kln_swa)
Dholuo-Kiswahili (luo_swa)
Each corpus contains approximately 30,000 sentence pairs. The dataset was created for use in training machine translation models, enabling translation from Kiswahili (the national language of Kenya) into indigenous… See the full description on the dataset page: https://huggingface.co/datasets/thinkKenya/kenyan-low-resource-language-data.multimodal_low-resource_language_translation
Dataset Card for Multimodal Low-Resource Language Translation Dataset
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is the dataset for our paper "From Text to Multi-Modal: Advancing Low-Resource-Language Translation through Synthetic Data Generation and Cross-Modal Alignments" accepted by the workshop LoResMT 2025 of NAACL 2025… See the full description on the dataset page: https://huggingface.co/datasets/qianstats/multimodal_low-resource_language_translation.translate-low-resource
Translation Dataset - Low Resource Indian Languages
Parallel translation datasets for 50 Indian languages, generated using GPT-5-mini for NLLB-200 finetuning.
Dataset Details
Total configs: 239
Examples per config: ~4,000
Total examples: ~956,000
Languages: 50 Indian languages across Indo-Aryan, Dravidian, Austroasiatic, and Sino-Tibetan families
Hub languages: English, Hindi, Bengali, Tamil, Odia, Assamese
Usage
Language Codes
Code
Language… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/translate-low-resource.lowresource
Dataset Card for "lowresource"
More Information needed
Adaption-low-resource-audio
Adaption Low-Resource Audio
A low-resource-language subset of
Reubencf/PolyglotAudio,
remastered with Adaption's Adaptive Data
platform. Each row carries the original Tatoeba-derived audio clip
alongside sharpened enhanced_prompt / enhanced_completion columns
so the data is ready for speech-model fine-tuning and evaluation on
languages that are typically under-represented in open ASR/TTS corpora.
Dataset size
3,704 rows of paired audio + text, spanning 10 languages… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-audio.low_resource_parallel_corpora
The Little Prince — multiparallel corpus (22 languages, RU pivot)
A sentence-level multiparallel corpus of Antoine de Saint-Exupéry's The Little Prince,
built around the classic Russian translation by Nora Gal as the pivot and covering
21 further editions, most of them in low-resource minority languages of Russia.
Every one of the 1,565 pivot sentences has exactly one aligned sentence in every
included language — a perfect N-way alignment (no gaps, no merges). The editions were… See the full description on the dataset page: https://huggingface.co/datasets/averoo/low_resource_parallel_corpora.low_resource_multilingual_sftlow_resource_multilingual_sft_short2low_resource_language
Beyond Log Likelihood Low Resource Language
This dataset bundle contains the low-resource language training/validation
parquet files and the MMLU-ProX-style multilingual multiple-choice test JSON
used by the Beyond-Log-Likelihood repository.
Low-Resource_Audio数据展示选取四种语言:西班牙语(es)与德语(de)作为基座语言,卡拜尔语(kab)与卢干达语(lg)作为新到来的语种,均取自 Common Voice Scripted Speech 26.0 的训练划分,每种语言各取其中一部分语句,音频统一重采样为 16 kHz 单声道。
这一组合有两点值得说明:es/de 与 kab/lg 分属不同语系,因此可以考察跨语系的迁移与干扰;而两侧在语料规模上存在明显落差,正好对应监督信息受限下的适应场景。
low_resource_multilinguallow-resource-naija-pretrainlow-resource-multilingual-doc-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
multilingual_doc_qa
This dataset contains over 10,000 question-answer pairs derived from multilingual document pages, covering languages such as Italian, German, Chinese, Portuguese, and Japanese. Each sample includes the original OCR text, page metadata, and specific queries regarding dates, titles, entities, or content details found within the documents. The data is… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/low-resource-multilingual-doc-qa.low_resource_multilingual_sft_shortasr-lwazi-low-resourcedlowresource_brx_doi_mni
Dataset Card for "lowresource_brx_doi_mni"
More Information needed
LocalAI_Low_Resourceslowresource_6k_sample
Dataset Card for "lowresource_6k_sample"
More Information needed
