CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes4k downloads4mo agoHugging Face02bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes227 downloads9d agoHugging Face03ImruQays /Rasaif-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period. Content Details Contained within this dataset are English translations of the following texts, sourced from the Rasaif website: A Muslim Manual of War Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes214 downloads3y agoHugging Face04ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K1 likes188 downloads2y agoHugging Face05shenasa /English-Persian-Parallel-Dataset English-Persian Parallel Dataset This repository provides access to a high-quality parallel dataset for English-to-Persian translation. The dataset has been curated for research purposes and is suitable for training and evaluating Neural Machine Translation (NMT) models. Download Link You can download the dataset using the following link: Download English-Persian Parallel Dataset Description The dataset contains aligned sentence pairs in English and Persian… See the full description on the dataset page: https://huggingface.co/datasets/shenasa/English-Persian-Parallel-Dataset.text1M<n<10M11 likes170 downloads1y agoHugging Face06Ghana-NLP /ENGLISH_TWI_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.text1K<n<10K3 likes143 downloads10mo agoHugging Face07Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes132 downloads4y agoHugging Face08Funghang /plasma-parallel-dbd-air Non-thermal Plasma Parallel DBD Air Dataset Overview This dataset contains experimental time-series measurements from a parallel Dielectric Barrier Discharge (DBD) plasma system in air at NTP. The dataset was collected using a digital oscilloscope and includes current-voltage waveforms measurements for plasma discharge characterization. Data Acquisition The experiments were conducted in the Physics Laboratory, Department of Physics, Kathmandu… See the full description on the dataset page: https://huggingface.co/datasets/Funghang/plasma-parallel-dbd-air.textfeature-extraction10K<n<100K3 likes126 downloads3mo agoHugging Face09israel /flores-paralleltabular1K<n<10K0 likes114 downloads2y agoHugging Face10Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.0. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K0 likes100 downloads18h agoHugging Face11Omarrran /kashmiri_English_parallel_corpus_49Kgated license: apache-2.0 task_categories: translation language: ks Usage Terms for this Dataset Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications. Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following: @misc {haq_nawaz_malik_2024, author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.text10K<n<100K2 likes84 downloads1y agoHugging Face12jaio98 /ParallelXNLIvartext10K<n<100K0 likes79 downloads5d agoHugging Face13Ghana-NLP /FANTE_ENGLISH_PARALLEL_TEXT GhanaNLP Fante and English Data Fante_to_English • 1 MB • XLS The GhanaNLP Fante dataset contains sentence pairs in Fante and English, designed to support translation models between these two languages. Fante is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/FANTE_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes72 downloads10mo agoHugging Face14DigitalUmuganda /NMT_Rwandan-Gazette_parallel_data_en_kin Dataset Details Dataset Description This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix Curated by: Digital Umuganda Language(s) (NLP): Kinyarwanda and English License: cc-by-4.0 Dataset Sources [optional] The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.texttranslation100K<n<1M3 likes68 downloads3y agoHugging Face15sello-ralethe /SA-Parallel-Corpora SA-Parallel-Corpora Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi bitext, drawn from South African government publications. Produced for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Code at https://github.com/sello-ralethe/SA-knowledge Structure One configuration per language pair, each with train, validation and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.tabular10K<n<100K0 likes67 downloads14d agoHugging Face16rakhine-nlp /rakhine-english-parallel-corpus 🌐 Rakhine–English Parallel Corpus A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP). 🎯 Purpose This dataset is designed to support: Machine Translation (MT) Neural Machine Translation (NMT) Language Modeling Low-resource NLP research Linguistic and dialect studies Language preservation and documentation 📌 Overview Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.texttranslationn<1K0 likes64 downloads3mo agoHugging Face17PrinceAlhassanNasamu /kusaal-english-parallel-corpus Kusaal-English Parallel Corpus The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset. This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.texttranslation10K<n<100K2 likes63 downloads2mo agoHugging Face18Ghana-NLP /TWI_ENGLISH_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/TWI_ENGLISH_PARALLEL_TEXT.text1K<n<10K1 likes57 downloads10mo agoHugging Face19abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face20Murtazali /quran-dargwa-parallel Quran Arabic–Dargwa Parallel Corpus A verse-aligned parallel corpus of the Quran in Arabic and Dargwa. The dataset contains 6,236 aligned records covering all 114 surahs. Each record contains an Arabic verse and its Dargwa translation. The Dargwa text is based on the translation by Magomed Gamidov, published by Yupiter in Makhachkala in 1995. The printed edition was digitized using OCR, corrected semi-automatically, and partially reviewed manually. A small number of OCR… See the full description on the dataset page: https://huggingface.co/datasets/Murtazali/quran-dargwa-parallel.tabulartranslation1K<n<10K0 likes57 downloads17d agoHugging Face21HiTZ /Parallel-XNLIvarBrief dataset description: Native: the native partition of XNLIeu (Heredia et al., 2024) adapted into three Basque dialects. Test: the test partition of XNLI adapted into three Basque dialects. All_dialects _together is a train/dev/test split that includes both native and test instances, stratified according to dialects. text10K<n<100K0 likes56 downloads3mo agoHugging Face22Horeknad /komi-russian-parallel-corpora Source Datasets 1 - news from the website of the Komi administration (https://rkomi.ru/) 2 - Komi media library (http://videocorpora.ru/) 3 - Millet porridge by Ivan Toropov (adaptation) Authors Shilova Nadezhda Chernousov Georgy texttranslation10K<n<100K2 likes55 downloads3y agoHugging Face23Ghana-NLP /EWE_ENGLISH_PARALLEL_TEXT GhanaNLP Ewe and English Data Ewe_to_English • 1 MB • XLS The GhanaNLP Ewe dataset contains sentence pairs in Ewe and English, designed to support translation models between these two languages. Ewe is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/EWE_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes54 downloads10mo agoHugging Face24QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes51 downloads2y agoHugging Face25Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes50 downloads3y agoHugging Face26navinaananthan /Kurdish-Sorani-Parallel-Corpustext100K<n<1M5 likes48 downloads3y agoHugging Face272ADT-Consulting /susu-parallel Susu (Soussou) Parallel and Monolingual Corpus A multi-source corpus for Susu (Soussou; ISO 639-3 sus), a Mande language of Guinea that is absent from NLLB-200 and from commercial MT systems. Built to train 2ADT-Consulting/nllb-susu-v2, one of the first open neural MT systems for Susu. Configurations Config Split #rows Columns sus-fr train / validation / test 114,503 / 1,000 / 1,000 sus, fr sus-en train / validation / test 111,013 / 991 / 992 sus, en… See the full description on the dataset page: https://huggingface.co/datasets/2ADT-Consulting/susu-parallel.tabulartranslation100K<n<1M1 likes48 downloads2mo agoHugging Face28DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes47 downloads4mo agoHugging Face29fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes46 downloads21d agoHugging Face30HackHedron /English_Telugu_Parallel_Corpustexttranslation100K<n<1M1 likes42 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.