datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english_karakalpak_parallel_corpus_v5
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.african-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.kashmiri_English_parallel_corpus_49K
license: apache-2.0
task_categories:
translation
language:
ks
Usage Terms for this Dataset
Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications.
Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following:
@misc {haq_nawaz_malik_2024,
author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.rakhine-english-parallel-corpus
🌐 Rakhine–English Parallel Corpus
A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP).
🎯 Purpose
This dataset is designed to support:
Machine Translation (MT)
Neural Machine Translation (NMT)
Language Modeling
Low-resource NLP research
Linguistic and dialect studies
Language preservation and documentation
📌 Overview
Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.kusaal-english-parallel-corpus
Kusaal-English Parallel Corpus
The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset.
This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.JParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.Kurdish-Sorani-Parallel-CorpusMyanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.English_Telugu_Parallel_CorpusCA-EN_Parallel_Corpus
Dataset Card for CA-EN Parallel Corpus
Dataset Description
Dataset Summary
The CA-EN Parallel Corpus is a Catalan-English dataset of parallel sentences created to
support Catalan in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between English and Catalan in any direction,
as well as Multilingual Machine Translation models.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EN_Parallel_Corpus.english_karakalpak_parallel_corpus_v10
English-Karakalpak Parallel Corpus v10.0
Dataset Description
English-Karakalpak Parallel Corpus v10.0 is a high-quality, finalized parallel dataset containing over 50,667 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems.
Language(s): English (en), Karakalpak (kaa)
Format: CSV (Comma-Separated Values)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v10.wolof-arabic-parallel-corpus
MudawanSn: A Gold-Standard Wolof--Arabic Parallel Corpus for Machine Translation
A publicly available parallel corpus for the Wolof–Arabic language pair, a gold-standard resource containing 1,271 sentence-aligned pairs. The corpus consists of manual translations from Wolof into Modern Standard Arabic (MSA). The source texts are drawn from the MasakhaNER corpus, covering politics, society, religion, and sports in Senegalese news discourse.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/mbaye930/wolof-arabic-parallel-corpus.French_Wolof_Various_Parallel_CorpusNepTam-A-Nepali-Tamang-Parallel-Corpus
🧾 NepTam — A Nepali–Tamang Parallel Corpus
Dataset Summary
NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains:
20K gold-standard human-translated sentence pairs, and
80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus.
Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.Deshika-Maharashtri_Prakrit_to_English_Parallel_CorpusMaharashtri Prakrit to English Parallel Corpus
Dataset Summary
This dataset contains parallel text data for translating from Maharashtri Prakrit (an ancient Indo-Aryan language) to English. It is designed to aid in developing machine translation systems, language models, and linguistic research for this underrepresented language. The dataset is collected from historical texts, scriptures, and scholarly resources.
Key Features:
Source Language: Maharashtri Prakrit
Target Language: English… See the full description on the dataset page: https://huggingface.co/datasets/VIITPune/Deshika-Maharashtri_Prakrit_to_English_Parallel_Corpus.english_karakalpak_parallel_corpus_v8
English-Karakalpak Parallel Corpus v8.0
Dataset Description
English-Karakalpak Parallel Corpus v8.0 is a high-quality, finalized parallel dataset containing over 37,257 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems.
Language(s): English (en), Karakalpak (kaa)
Format: CSV (Comma-Separated Values)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v8.A-Wolof-Arabic-Parallel-Corpusbhasaflow-khasi-english-parallel-corpus-v1
BhasaFlow Khasi-English Parallel Corpus v1
By Medharvix Systems Private Limited
Overview
A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India.
Dataset Structure
Column
Description
sentence_id
Unique sentence identifier
english_text
English sentence
khasi_text
Khasi translation
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.ted_talks_multilingual_parallel_corpusmultilingual_parallel_corpusenglish_karakalpak_parallel_corpus_v3-4
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus v3-4 is a high-quality dataset containing 2,722 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v3-4.english_karakalpak_parallel_corpus_v7
English-Karakalpak Parallel Corpus v7.0
Dataset Description
English-Karakalpak Parallel Corpus v7.0 is a high-quality, finalized parallel dataset containing over 32,972 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems.
Language(s): English (en), Karakalpak (kaa)
Format: CSV (Comma-Separated Values)
License: MIT
Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v7.english_karakalpak_parallel_corpus_v9
English-Karakalpak Parallel Corpus v9.0
Dataset Description
English-Karakalpak Parallel Corpus v9.0 is a high-quality, finalized parallel dataset containing over 40,513 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems.
Language(s): English (en), Karakalpak (kaa)
Format: CSV (Comma-Separated Values)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v9.HFcourse-english-burmese-parallel-corpus
HFcourse-English-Burmese-Parallel-Corpus
Dataset Description
Dataset Summary
The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.pali-myanmar-parallel-corpus-1kenglish_karakalpak_parallel_corpus_v1
English-Karakalpak Parallel Corpus (en-kaa)
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.English_Telugu_Parallel_Corpustamajaq-english-parallel-corpus
Tawallammat Tamajaq - English Parallel Corpus (1.3K Sentences)
Dataset Description
This dataset is a parallel corpus containing translated sentence pairs between the Tawallammat Tamajaq (ttq) language and English (en). The data has been curated, filtered, and contributed from open-source platforms like Tatoeba and Glosbe by contributor Tamajiq1286.
The primary goal of this project is to support low-resource language development and provide an open digital… See the full description on the dataset page: https://huggingface.co/datasets/Tamajeq1286/tamajaq-english-parallel-corpus.
