datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
Egyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.cantonese-chinese-parallel-corpus
Dataset Summary
This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation.
The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation.
Languages
Cantonese (yue)
Simplified Chinese (zh)
Dataset Structure
Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.C-Rust-parallel-corpus
Dataset Card for C-to-Rust Parallel Semantic Similarity Corpus
Dataset Summary
The C-to-Rust Parallel Semantic Similarity Corpus is a curated dataset consisting of 1,886 aligned, function-level C and Rust code pairs. It was developed to evaluate cross-language semantic similarity and functional equivalence between a traditional legacy language (C) and a modern memory-safe language (Rust).
The source code snippets are drawn from accepted competitive programming… See the full description on the dataset page: https://huggingface.co/datasets/hejlevoj/C-Rust-parallel-corpus.icelandic-parallel-abstracts-corpus-IPACSee https://arxiv.org/abs/2108.05289
english-manipuri-parallel-corpus
English-Manipuri Corpus
About
This English-Manipuri corpus contains an expanded parallel corpus for English-Manipuri of the following paper.
Dataset Statistics
Bible Dataset contains approx. 31K parallel sentences.
PIB-PMI Dataset contains approx. 500K parallel sentences.
How to Use
You can load the dataset using the Huggingface datasets library:
from datasets import load_dataset
# Load the bible dataset
bible_dataset = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/joyson117/english-manipuri-parallel-corpus.cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
hindi-kumaoni-parallel-corpus
Hindi ↔ Kumaoni Parallel Corpus
First publicly available clean Hindi-Kumaoni parallel translation dataset.
Dataset Details
Language pair: Hindi (hi) ↔ Kumaoni (kum, ISO 639-3: kfy)
Size: 920 pairs (736 train / 92 dev / 92 test)
Script: Devanagari
Dialect: Central Kumaoni (Almora)
Sources
speakkumaoni.com (structured lessons)
euttaranchal.com (language lessons)
Wikipedia EN Kumaoni language page
Format
Each JSONL line:
{"translation": {"hi": "..."… See the full description on the dataset page: https://huggingface.co/datasets/swapedoc/hindi-kumaoni-parallel-corpus.english_pashto_parallel-corpus
English–Pashto Parallel Corpus
A high-quality English–Pashto parallel corpus designed for machine translation, supervised fine-tuning, Pashto LLM training, and corpus quality research.
📁 Dataset Structure
1. translation_clean.jsonl
High-quality parallel data suitable for:
Machine Translation
Supervised Fine-Tuning (SFT)
Pashto LLM Training
Instruction Tuning
Reasoning Alignment
Each record follows:
{
"id": 12345,
"source": "English… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_pashto_parallel-corpus.Chinese-English-Parallel-Synonym-Corpus-75kaz-eng-parallel-corpusparallel_corpus_russian_rsl_glossesEnglish_Nuer_parallel_translation_open_corpusChinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat
Chinese-English Parallel Translation Corpus (Chinese Source Text & English Translation)
A Chinese–English parallel corpus resource for translation and cross-lingual alignment applications, providing one-to-one bilingual text pairs: Chinese source texts aligned with their corresponding English translations. The data covers common writing styles and domains, making it suitable for parallel alignment, translation modeling, and cross-lingual representation learning.
It supports… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat.en-ja-parallel-corpus-augmentedluganda-english-parallel-corpus
English-Luganda Parallel Corpus for Translation
Dataset Description
This dataset contains parallel sentences in English (en) and Luganda (lg), designed primarily for training and fine-tuning machine translation models. The data consists of sentence pairs extracted from a source document.
Languages
English (en)
Luganda (lg) - ISO 639-1 code: lg
Data Format
The dataset is provided in a format compatible with the Hugging Face datasets library. Each… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-parallel-corpus.JParaCrawl-Filtered-English-Japanese-Parallel-Corpus-textenglish-nuer-dinka-parallel-corpus
English–Nuer–Dinka Parallel Corpus
Overview
The English–Nuer–Dinka Parallel Corpus is a multilingual parallel dataset created to support research on low-resource African languages. The corpus contains aligned text in English, Nuer, and Dinka for use in Natural Language Processing (NLP), Machine Translation (MT), multilingual language modeling, and language preservation.
The primary goal of this project is to increase the digital presence of Nuer and Dinka while… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/english-nuer-dinka-parallel-corpus.english_nuer-thok_naath_conversational_parallel_corpus
English–Nuer (Thok Naath) Conversational Parallel Corpus
Overview
The English–Nuer (Thok Naath) Conversational Parallel Corpus is a bilingual dataset consisting of aligned conversational sentence pairs in English and Nuer (Thok Naath).
This dataset was created to support research on low-resource African languages, with a particular focus on conversational AI, machine translation, multilingual language models, and language preservation.
As one of the few publicly… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/english_nuer-thok_naath_conversational_parallel_corpus.parallel-corpus_en-amateso-english-parallel-corpusodia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.zu-en_parallel-corpus_xsm
