datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.english_karakalpak_parallel_corpus_v5
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.Rasaif-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period.
Content Details
Contained within this dataset are English translations of the following texts, sourced from the Rasaif website:
A Muslim Manual of War
Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.English-Persian-Parallel-Dataset
English-Persian Parallel Dataset
This repository provides access to a high-quality parallel dataset for English-to-Persian translation. The dataset has been curated for research purposes and is suitable for training and evaluating Neural Machine Translation (NMT) models.
Download Link
You can download the dataset using the following link:
Download English-Persian Parallel Dataset
Description
The dataset contains aligned sentence pairs in English and Persian… See the full description on the dataset page: https://huggingface.co/datasets/shenasa/English-Persian-Parallel-Dataset.ENGLISH_TWI_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.korean-parallel-corporaplasma-parallel-dbd-air
Non-thermal Plasma Parallel DBD Air Dataset
Overview
This dataset contains experimental time-series measurements from a parallel Dielectric Barrier Discharge (DBD) plasma system in air at NTP.
The dataset was collected using a digital oscilloscope and includes current-voltage waveforms measurements for plasma discharge characterization.
Data Acquisition
The experiments were conducted in the Physics Laboratory, Department of Physics, Kathmandu… See the full description on the dataset page: https://huggingface.co/datasets/Funghang/plasma-parallel-dbd-air.flores-parallelafrican-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.kashmiri_English_parallel_corpus_49K
license: apache-2.0
task_categories:
translation
language:
ks
Usage Terms for this Dataset
Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications.
Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following:
@misc {haq_nawaz_malik_2024,
author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.ParallelXNLIvarFANTE_ENGLISH_PARALLEL_TEXT
GhanaNLP Fante and English
Data
Fante_to_English
• 1 MB • XLS
The GhanaNLP Fante dataset contains sentence pairs in Fante and English,
designed to support translation models between these two languages.
Fante is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/FANTE_ENGLISH_PARALLEL_TEXT.NMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.SA-Parallel-Corpora
SA-Parallel-Corpora
Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi
bitext, drawn from South African government publications.
Produced for the doctoral thesis Injecting Commonsense Knowledge into
Pretrained Language Models for Low Resource Languages (University of Cape Town,
2026). Code at https://github.com/sello-ralethe/SA-knowledge
Structure
One configuration per language pair, each with train, validation
and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.rakhine-english-parallel-corpus
🌐 Rakhine–English Parallel Corpus
A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP).
🎯 Purpose
This dataset is designed to support:
Machine Translation (MT)
Neural Machine Translation (NMT)
Language Modeling
Low-resource NLP research
Linguistic and dialect studies
Language preservation and documentation
📌 Overview
Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.kusaal-english-parallel-corpus
Kusaal-English Parallel Corpus
The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset.
This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.TWI_ENGLISH_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/TWI_ENGLISH_PARALLEL_TEXT.Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.quran-dargwa-parallel
Quran Arabic–Dargwa Parallel Corpus
A verse-aligned parallel corpus of the Quran in Arabic and Dargwa.
The dataset contains 6,236 aligned records covering all 114 surahs. Each record contains an Arabic verse and its Dargwa translation.
The Dargwa text is based on the translation by Magomed Gamidov, published by Yupiter in Makhachkala in 1995. The printed edition was digitized using OCR, corrected semi-automatically, and partially reviewed manually. A small number of OCR… See the full description on the dataset page: https://huggingface.co/datasets/Murtazali/quran-dargwa-parallel.Parallel-XNLIvarBrief dataset description:
Native: the native partition of XNLIeu (Heredia et al., 2024) adapted into three Basque dialects.
Test: the test partition of XNLI adapted into three Basque dialects.
All_dialects _together is a train/dev/test split that includes both native and test instances, stratified according to dialects.
komi-russian-parallel-corpora
Source Datasets
1 - news from the website of the Komi administration (https://rkomi.ru/)
2 - Komi media library (http://videocorpora.ru/)
3 - Millet porridge by Ivan Toropov (adaptation)
Authors
Shilova Nadezhda
Chernousov Georgy
EWE_ENGLISH_PARALLEL_TEXT
GhanaNLP Ewe and English
Data
Ewe_to_English
• 1 MB • XLS
The GhanaNLP Ewe dataset contains sentence pairs in Ewe and English,
designed to support translation models between these two languages.
Ewe is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/EWE_ENGLISH_PARALLEL_TEXT.crh-parallel-corpora-document-level-noisyJParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.Kurdish-Sorani-Parallel-Corpussusu-parallel
Susu (Soussou) Parallel and Monolingual Corpus
A multi-source corpus for Susu (Soussou; ISO 639-3 sus), a Mande language of
Guinea that is absent from NLLB-200 and from commercial MT systems. Built to train
2ADT-Consulting/nllb-susu-v2,
one of the first open neural MT systems for Susu.
Configurations
Config
Split
#rows
Columns
sus-fr
train / validation / test
114,503 / 1,000 / 1,000
sus, fr
sus-en
train / validation / test
111,013 / 991 / 992
sus, en… See the full description on the dataset page: https://huggingface.co/datasets/2ADT-Consulting/susu-parallel.Myanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.English_Telugu_Parallel_Corpus
