datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pontoon-Translations
Dataset Card for Pontoon Translations
This is a dataset containing strings from various Mozilla projects on Mozilla's Pontoon localization platform and their translations into more than 200 languages.
Source strings are in English.
To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this:
from datasets import load_dataset
dataset = load_dataset("ayymen/Pontoon-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Pontoon-Translations.ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.idiom_translation_finalJP-TH_Literary_Translation_URL_Alignment_Index
JP–TH Literary Translation URL Alignment Index
This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting.
Overview
The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.test-translation-datasetcross-species-translational-alignment
Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21
Goal: build a training substrate for detecting subtle / pre-histopathological
toxicity signatures in animal transcriptome data, with mechanism-of-toxicity
labels attached. This directory contains the compound-level linkage layer:
every compound that has rat in-vivo perturbation data cross-referenced to Tox21
mechanism assays via standardized chemical identifiers.
Background — the hackathon
Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.ES-VA_translation_test
Subtask (ES-VA_translation) of Phrases adaptability task
This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity.
Subsequently, the selected sentences were translated from… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/ES-VA_translation_test.ko-en-structured-translations
Korean–English Multistyle Parallel Corpus
한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스
소개(Introduction)
저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다.
기존 공개 데이터셋은
도메인 다양성이 부족하거나
문체가 한국 사용자에게 자연스럽지 않거나
전반적으로 문장의 퀄리티가 매우 부족하여
학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다.
이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다.
한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.Pontoon-Translations
Pontoon Translations
Amazigh subset of Pontoon Translations.
aihub-koen-translation-integrated-large-10m
AI Hub Ko-En Translation Dataset (Integrated)
AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다.
병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다.
base-10m: 병합 데이터 100% 사용, 총 10,416,509개
mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개
tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개
Subsets
활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다.
전문분야 한영 말뭉치 (111)
총 개수: 1,350,000
중복 제거 후 개수: 1,350,000
사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-large-10m.tagalog-filipino-english-translationThis dataset is a Tagalog-English translation data. It is a compiled comma-separated values dataset from different
existing HuggingFace and External dataset.
Here are the collected and compiled data:
saillab/alpaca_tamil_taco
DIBT/MPEP_FILIPINO
Nag, S., Ma, S., Ntalli, A., & Dulay, K. M. (2024, June 10). TalkTogether. https://doi.org/10.17605/OSF.IO/3ZDFN
basque_dialect_machine_translationswiss-legal-translation
Swiss Legal Translation Dataset
A large-scale parallel corpus of Swiss legal texts in German, French, and Italian, extracted from official government sources.
Dataset Overview
Metric
Value
Total articles
62,594
Languages
German, French, Italian
All 3 languages
28,000 articles (44.7%)
At least 2 languages
62,594 articles (100%)
Source laws
743+ unique laws
File size
~88 MB
Language Coverage by Source
Source
Articles
DE
FR
IT… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/swiss-legal-translation.urdu-idioms-with-english-translationWolof-to-French_Translation-Dataset
Dataset Wolof ↔ Français
🧩 Présentation
Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit :
Wolof (input)
Français (target)
Phrase en Wolof
Phrase correspondante en Français
Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français.
📚 Provenance et nettoyage
Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.tibetan-to-spanish-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the Spanish translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
Luganda_Sci-Math-Bio_Translations
Luganda Sci-Math-Bio Translations
This dataset contains Luganda and English translations of biologicial, mathematical and scientific terms
Cantonese_English_Translation
Cantonese_English_Translation
Overview | 總括
This dataset provides parallel text translations between Cantonese and English, suitable for research and development in natural language processing and machine translation. | 呢個資料庫提供廣東話同英文嘅對應翻譯,啱晒用嚟做語言處理同機器翻譯嘅研究同開發。
Dataset Structure | 資料組織
english_cantonese_translation.csv: Contains two fields: "english" and "cantonese". | 有兩個位: "english" 同 "cantonese"。
Usage Example | 用法例子
import pandas as pd
# Load… See the full description on the dataset page: https://huggingface.co/datasets/lordjia/Cantonese_English_Translation.massive_translation_dataset
Dataset Card for Massive Dataset for Translation
Dataset Summary
This dataset is derived from AmazonScience/MASSIVE dataset for translation task purpose.
Supported Tasks and Leaderboards
Translation
Languages
English (en_US)
German (de_DE)
Hindi (hi_IN)
Spanish (es_ES)
French (fr_FR)
Italian (it_IT)
Arabic (ar_SA)
Dutch (nl_NL)
Japanese (ja_JP)
Portugese (pt_PT)
Amazigh-Quran-Translation-Jouhadi
Dataset Card: Tamazight (Tifinagh) Quran Translation - Lahoucine Jouhadi
This dataset provides a digitized, partial translation of the meanings of the Holy Quran into Amazigh (Tachelhit) using the Neo-Tifinagh script. The content is based on the full translation work of Lahoucine Jouhadi (Lhocine Jouhadi Baamrani) based on Warsh recitation used in Morocco.
Original Sources & References
Author's Website - Down currently: Jouhadi Lahoussine Publications… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Quran-Translation-Jouhadi.English-to-Afar-language-translation
Author
Created by Charif Ayfarah.
Contact: afbarit@gmail.com
License
Licensed under CC BY 4.0. You are free to use, modify, and distribute this dataset, including for commercial purposes, as long as you give appropriate credit.
thai-local-language-translation-dataset
Thai Local Language Translation Dataset
Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus.
The dataset support Khummuang, Korat, and Pattani.
Reference
Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.aihub-koen-translation-integrated-small-100k
AI Hub Ko-En Translation Dataset (Integrated)
AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다.
병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다.
base-10m: 병합 데이터 100% 사용, 총 10,416,509개
mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개
tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개
Subsets
활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다.
전문분야 한영 말뭉치 (111)
총 개수: 1,350,000
중복 제거 후 개수: 1,350,000
사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-small-100k.aihub-koen-translation-integrated-base-1m
AI Hub Ko-En Translation Dataset (Integrated)
AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다.
병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다.
base-10m: 병합 데이터 100% 사용, 총 10,416,509개
mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개
tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개
Subsets
활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다.
전문분야 한영 말뭉치 (111)
총 개수: 1,350,000
중복 제거 후 개수: 1,350,000
사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-base-1m.ViKm-Translation-Task
ViKm-Trans
A high-quality synthetic Vietnamese–Khmer parallel corpus.
Overview
ViKm-Trans is a synthetic parallel corpus for Vietnamese ↔ Khmer machine translation.
Due to the scarcity of publicly available Vietnamese–Khmer parallel data, we propose a synthetic data generation framework that leverages abundant Vietnamese monolingual corpora together with large language models to construct high-quality parallel sentence pairs.
The dataset was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/ViKm-Translation-Task.aihub-kozh-translation-integrated-large-5.9m
AI Hub Ko-Zh Translation Dataset (Integrated)
AI Hub의 한-중 번역 관련 데이터셋 10개를 병합한 자료입니다. 병합 시 총 데이터 개수는 5,934,596개이며, 이중 10,000개의 validation set와 2,000개의 test set가 분리되어 모든 데이터 사이즈(large-5.9m, base-1m, small-100k)에서 동일하게 사용됩니다.
large-5.9m (train): 병합 데이터 100% 사용; 총 5,922,596개
base-1m (train): 병합 데이터 중 1M개 사용; 총 1,000,000개
small-100k (train): 병합 데이터 중 100K개 사용; 총 100,000개
Subsets
Name
Total Size
Chinese Size (Utilized Only)
URL
Datasetkey (AIHub)
한국어-중국어 번역… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-kozh-translation-integrated-large-5.9m.vietnamese-nom-poetry-translationsinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.cross_species_leaf_absolute_translationcarrot-engine-normalization-translation-v2
