CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aletheia-ng /low_resource_languages_pretrain_data5text100M<n<1B0 likes1.2k downloads11mo agoHugging Face02Aletheia-ng /low_resource_languages_pretrain_data2text100M<n<1B0 likes1.1k downloads1y agoHugging Face03BeardedMonster /low_resource_languages_pretrain_data8text100M<n<1B0 likes891 downloads9mo agoHugging Face04Aletheia-ng /low_resource_languages_pretraintext100M<n<1B1 likes699 downloads1y agoHugging Face05surrey-nlp /Low-resource-QE-DA-dataset Low-resource QE-DA Dataset Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE. Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv) Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.tabularother100K<n<1M0 likes601 downloads10mo agoHugging Face06Aletheia-ng /low_resource_languages_pretrain_data4text100M<n<1B0 likes498 downloads11mo agoHugging Face07Aletheia-ng /low_resource_languages_pretrain_datatext100M<n<1B0 likes450 downloads1y agoHugging Face08Pika4028 /low-resource-rag-indexes Low-Resource RAG: Wikipedia FAISS Indexes (BGE-M3) FAISS indexes over Wikipedia 2023 passages for four languages, embedded with BAAI/bge-m3. Source corpus: CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Each language folder has two files aligned row-by-row: index.faiss — FAISS IndexFlatIP, dim=1024, L2-normalized BGE-M3 vectors dataset/ — Arrow dataset with _id, title, text per passage Languages Language Code Passages Size Hindi hi TBD 2.6G… See the full description on the dataset page: https://huggingface.co/datasets/Pika4028/low-resource-rag-indexes.0 likes113 downloads4mo agoHugging Face09Subayyal /Urdu-Low-Resource-Language-Dataset Dataset Card for Low-Resource-Language-Dataset This dataset is designed to aid Natural Language Processing (NLP) research on low-resource languages, particularly Urdu. It includes structured datasets and preprocessing tools curated from the BBC Urdu website. Dataset Details Dataset Description This dataset contains articles, summaries, and topics scraped from BBC Urdu. It is structured into training and testing datasets for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Subayyal/Urdu-Low-Resource-Language-Dataset.textsummarization1 likes107 downloads16d agoHugging Face10rl-low-resource /rumantsch-varieties-sentencestext100K<n<1M0 likes96 downloads7mo agoHugging Face11limhyeonseok /mgsm-all-low-resource-translatedtext1K<n<10K0 likes86 downloads9mo agoHugging Face12Reubencf /Adaption-low-resource-doc-qa Adaption Low-Resource Document Q/A This dataset is a remastered version of Reubencf/magazines-multilingual-vqa prepared using Adaption's Adaptive Data platform, with a deliberate focus on low-resource source languages — the languages that are underrepresented in most open multimodal datasets. What's inside 10,200 rows of multilingual document question-answer pairs grounded in public-domain magazine / newspaper pages from archive.org. Every row carries verbatim OCR in… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-doc-qa.imagevisual-question-answering1K<n<10K0 likes75 downloads5mo agoHugging Face13thinkKenya /kenyan-low-resource-language-datagated Low-Resource Language Data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin, and Dholuo Description This dataset consists of three parallel corpora: Kidaw'ida (Dawida)-Kiswahili (dav_swa) Kalenjin-Kiswahili (kln_swa) Dholuo-Kiswahili (luo_swa) Each corpus contains approximately 30,000 sentence pairs. The dataset was created for use in training machine translation models, enabling translation from Kiswahili (the national language of Kenya) into indigenous… See the full description on the dataset page: https://huggingface.co/datasets/thinkKenya/kenyan-low-resource-language-data.texttranslation10K<n<100K7 likes60 downloads2y agoHugging Face14qianstats /multimodal_low-resource_language_translation Dataset Card for Multimodal Low-Resource Language Translation Dataset This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This is the dataset for our paper "From Text to Multi-Modal: Advancing Low-Resource-Language Translation through Synthetic Data Generation and Cross-Modal Alignments" accepted by the workshop LoResMT 2025 of NAACL 2025… See the full description on the dataset page: https://huggingface.co/datasets/qianstats/multimodal_low-resource_language_translation.translation1 likes59 downloads1y agoHugging Face15ayush-shunyalabs /translate-low-resource Translation Dataset - Low Resource Indian Languages Parallel translation datasets for 50 Indian languages, generated using GPT-5-mini for NLLB-200 finetuning. Dataset Details Total configs: 239 Examples per config: ~4,000 Total examples: ~956,000 Languages: 50 Indian languages across Indo-Aryan, Dravidian, Austroasiatic, and Sino-Tibetan families Hub languages: English, Hindi, Bengali, Tamil, Odia, Assamese Usage Language Codes Code Language… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/translate-low-resource.text1M<n<10M0 likes56 downloads7mo agoHugging Face16akashmadisetty /lowresource Dataset Card for "lowresource" More Information needed text10K<n<100K0 likes35 downloads1y agoHugging Face17Reubencf /Adaption-low-resource-audio Adaption Low-Resource Audio A low-resource-language subset of Reubencf/PolyglotAudio, remastered with Adaption's Adaptive Data platform. Each row carries the original Tatoeba-derived audio clip alongside sharpened enhanced_prompt / enhanced_completion columns so the data is ready for speech-model fine-tuning and evaluation on languages that are typically under-represented in open ASR/TTS corpora. Dataset size 3,704 rows of paired audio + text, spanning 10 languages… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-audio.audioautomatic-speech-recognition1K<n<10K1 likes35 downloads5mo agoHugging Face18averoo /low_resource_parallel_corpora The Little Prince — multiparallel corpus (22 languages, RU pivot) A sentence-level multiparallel corpus of Antoine de Saint-Exupéry's The Little Prince, built around the classic Russian translation by Nora Gal as the pivot and covering 21 further editions, most of them in low-resource minority languages of Russia. Every one of the 1,565 pivot sentences has exactly one aligned sentence in every included language — a perfect N-way alignment (no gaps, no merges). The editions were… See the full description on the dataset page: https://huggingface.co/datasets/averoo/low_resource_parallel_corpora.tabulartranslation1K<n<10K9 likes35 downloads2mo agoHugging Face19ccibeekeoc42 /low_resource_multilingual_sfttext10K<n<100K0 likes29 downloads2y agoHugging Face20ccibeekeoc42 /low_resource_multilingual_sft_short2text10K<n<100K1 likes27 downloads2y agoHugging Face21gaotang /low_resource_language Beyond Log Likelihood Low Resource Language This dataset bundle contains the low-resource language training/validation parquet files and the MMLU-ProX-style multilingual multiple-choice test JSON used by the Beyond-Log-Likelihood repository. text-generation0 likes22 downloads4mo agoHugging Face22dasdfdaa /Low-Resource_Audio数据展示选取四种语言:西班牙语(es)与德语(de)作为基座语言,卡拜尔语(kab)与卢干达语(lg)作为新到来的语种,均取自 Common Voice Scripted Speech 26.0 的训练划分,每种语言各取其中一部分语句,音频统一重采样为 16 kHz 单声道。 这一组合有两点值得说明:es/de 与 kab/lg 分属不同语系,因此可以考察跨语系的迁移与干扰;而两侧在语料规模上存在明显落差,正好对应监督信息受限下的适应场景。 0 likes21 downloads2d agoHugging Face23ccibeekeoc42 /low_resource_multilingualtext10K<n<100K0 likes18 downloads2y agoHugging Face24BeardedMonster /low-resource-naija-pretraintext100K<n<1M0 likes15 downloads10mo agoHugging Face25Reubencf /low-resource-multilingual-doc-qa This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. multilingual_doc_qa This dataset contains over 10,000 question-answer pairs derived from multilingual document pages, covering languages such as Italian, German, Chinese, Portuguese, and Japanese. Each sample includes the original OCR text, page metadata, and specific queries regarding dates, titles, entities, or content details found within the documents. The data is… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/low-resource-multilingual-doc-qa.text1K<n<10K0 likes15 downloads5mo agoHugging Face26ccibeekeoc42 /low_resource_multilingual_sft_shorttext10K<n<100K0 likes13 downloads2y agoHugging Face27zionia /asr-lwazi-low-resourced0 likes13 downloads1y agoHugging Face28akashmadisetty /lowresource_brx_doi_mni Dataset Card for "lowresource_brx_doi_mni" More Information needed text10K<n<100K0 likes11 downloads1y agoHugging Face29Ender68 /LocalAI_Low_Resourcestext-generation1B<n<10B0 likes9 downloads11mo agoHugging Face30akashmadisetty /lowresource_6k_sample Dataset Card for "lowresource_6k_sample" More Information needed text1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.