CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes227 downloads8d agoHugging Face02Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.0. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K0 likes100 downloads11h agoHugging Face03Omarrran /kashmiri_English_parallel_corpus_49Kgated license: apache-2.0 task_categories: translation language: ks Usage Terms for this Dataset Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications. Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following: @misc {haq_nawaz_malik_2024, author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.text10K<n<100K2 likes84 downloads1y agoHugging Face04rakhine-nlp /rakhine-english-parallel-corpus 🌐 Rakhine–English Parallel Corpus A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP). 🎯 Purpose This dataset is designed to support: Machine Translation (MT) Neural Machine Translation (NMT) Language Modeling Low-resource NLP research Linguistic and dialect studies Language preservation and documentation 📌 Overview Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.texttranslationn<1K0 likes64 downloads3mo agoHugging Face05PrinceAlhassanNasamu /kusaal-english-parallel-corpus Kusaal-English Parallel Corpus The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset. This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.texttranslation10K<n<100K2 likes63 downloads2mo agoHugging Face06abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face07Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes50 downloads3y agoHugging Face08navinaananthan /Kurdish-Sorani-Parallel-Corpustext100K<n<1M5 likes48 downloads3y agoHugging Face09DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes47 downloads4mo agoHugging Face10fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes46 downloads21d agoHugging Face11HackHedron /English_Telugu_Parallel_Corpustexttranslation100K<n<1M1 likes42 downloads1y agoHugging Face12projecte-aina /CA-EN_Parallel_Corpus Dataset Card for CA-EN Parallel Corpus Dataset Description Dataset Summary The CA-EN Parallel Corpus is a Catalan-English dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between English and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EN_Parallel_Corpus.tabulartranslation10M<n<100M1 likes40 downloads1y agoHugging Face13bekan /english_karakalpak_parallel_corpus_v10 English-Karakalpak Parallel Corpus v10.0 Dataset Description English-Karakalpak Parallel Corpus v10.0 is a high-quality, finalized parallel dataset containing over 50,667 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v10.texttranslation10K<n<100K0 likes39 downloads8d agoHugging Face14mbaye930 /wolof-arabic-parallel-corpus MudawanSn: A Gold-Standard Wolof--Arabic Parallel Corpus for Machine Translation A publicly available parallel corpus for the Wolof–Arabic language pair, a gold-standard resource containing 1,271 sentence-aligned pairs. The corpus consists of manual translations from Wolof into Modern Standard Arabic (MSA). The source texts are drawn from the MasakhaNER corpus, covering politics, society, religion, and sports in Senegalese news discourse. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/mbaye930/wolof-arabic-parallel-corpus.texttranslation1K<n<10K3 likes37 downloads3mo agoHugging Face15Oumar199 /French_Wolof_Various_Parallel_Corpustexttranslation1K<n<10K2 likes36 downloads2y agoHugging Face16ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K0 likes29 downloads11mo agoHugging Face17VIITPune /Deshika-Maharashtri_Prakrit_to_English_Parallel_CorpusMaharashtri Prakrit to English Parallel Corpus Dataset Summary This dataset contains parallel text data for translating from Maharashtri Prakrit (an ancient Indo-Aryan language) to English. It is designed to aid in developing machine translation systems, language models, and linguistic research for this underrepresented language. The dataset is collected from historical texts, scriptures, and scholarly resources. Key Features: Source Language: Maharashtri Prakrit Target Language: English… See the full description on the dataset page: https://huggingface.co/datasets/VIITPune/Deshika-Maharashtri_Prakrit_to_English_Parallel_Corpus.texttranslation1K<n<10K1 likes26 downloads2y agoHugging Face18bekan /english_karakalpak_parallel_corpus_v8 English-Karakalpak Parallel Corpus v8.0 Dataset Description English-Karakalpak Parallel Corpus v8.0 is a high-quality, finalized parallel dataset containing over 37,257 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v8.texttranslation10K<n<100K1 likes25 downloads2mo agoHugging Face19geekdiop /A-Wolof-Arabic-Parallel-Corpustext1K<n<10K1 likes19 downloads7d agoHugging Face20MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-corpus-v1 BhasaFlow Khasi-English Parallel Corpus v1 By Medharvix Systems Private Limited Overview A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India. Dataset Structure Column Description sentence_id Unique sentence identifier english_text English sentence khasi_text Khasi translation Usage from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.texttranslationn<1K19 likes19 downloads5mo agoHugging Face21TankuVie /ted_talks_multilingual_parallel_corpustext10K<n<100K2 likes18 downloads3y agoHugging Face22tachiwin /multilingual_parallel_corpustext10K<n<100K0 likes18 downloads9d agoHugging Face23bekan /english_karakalpak_parallel_corpus_v3-4 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus v3-4 is a high-quality dataset containing 2,722 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v3-4.texttranslation1K<n<10K0 likes18 downloads10mo agoHugging Face24bekan /english_karakalpak_parallel_corpus_v7 English-Karakalpak Parallel Corpus v7.0 Dataset Description English-Karakalpak Parallel Corpus v7.0 is a high-quality, finalized parallel dataset containing over 32,972 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v7.texttranslation10K<n<100K0 likes15 downloads6mo agoHugging Face25bekan /english_karakalpak_parallel_corpus_v9 English-Karakalpak Parallel Corpus v9.0 Dataset Description English-Karakalpak Parallel Corpus v9.0 is a high-quality, finalized parallel dataset containing over 40,513 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v9.texttranslation10K<n<100K0 likes14 downloads2mo agoHugging Face26kalixlouiis /HFcourse-english-burmese-parallel-corpus HFcourse-English-Burmese-Parallel-Corpus Dataset Description Dataset Summary The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.texttranslation1K<n<10K11 likes13 downloads5mo agoHugging Face27kalixlouiis /pali-myanmar-parallel-corpus-1ktexttranslation1K<n<10K5 likes13 downloads5mo agoHugging Face28bekan /english_karakalpak_parallel_corpus_v1 English-Karakalpak Parallel Corpus (en-kaa) Dataset Description English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.texttranslation10K<n<100K2 likes12 downloads10mo agoHugging Face29RushikaSritha11 /English_Telugu_Parallel_Corpustexttranslation100K<n<1M0 likes12 downloads10mo agoHugging Face30Tamajeq1286 /tamajaq-english-parallel-corpus Tawallammat Tamajaq - English Parallel Corpus (1.3K Sentences) Dataset Description This dataset is a parallel corpus containing translated sentence pairs between the Tawallammat Tamajaq (ttq) language and English (en). The data has been curated, filtered, and contributed from open-source platforms like Tatoeba and Glosbe by contributor Tamajiq1286. The primary goal of this project is to support low-resource language development and provide an open digital… See the full description on the dataset page: https://huggingface.co/datasets/Tamajeq1286/tamajaq-english-parallel-corpus.text1K<n<10K3 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.