datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.norsumm-nob-nno-translation
Nynorsk-Bokmål translation pairs
A multi-sentence parallel corpus of manual Nynorsk-Bokmål translations. These translations were extracted from the SamiaT/NorSumm dataset. You can read more about how the original dataset was created (including details about the manual translation process) in Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles by Samia Touileb et al..
Contact
David Samuel (davisamu@ifi.uio.no)… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norsumm-nob-nno-translation.speech-translation-and-summarization
English-Centric Multilingual Audio Dataset
This dataset contains generated article and summary audio for English-centric multilingual directions.
Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits.
Included directions
amharic_english / english_amharic
arabic_english / english_arabic
bengali_english / english_bengali
chinese_simplified_english / english_chinese_simplified
english_english
french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.Europarl-Translation-Instruct
Dataset Card for Europarl-Translation-Instruct
Waifu to catch your attention.
Dataset Details
Dataset Description
europarl-translation-instruct is a translation instruct dataset built from europarl data.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): English instruct (but various languages in)
License: cc-by-sa-4.0
Dataset Sources
Source Data: https://www.statmt.org/europarl/ (Transcript source)
Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.BenchMAX_General_Translation
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain.
We collect parallel test data from Flore-200, TED-talk, and WMT24.
Usage
Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.BenchMAX_Domain_Translation
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Domain_Translation is a dataset of BenchMAX, which evaluates the translation capability on specific domains.
We collect the domain multi-way parallel data from other tasks in BenchMAX, such as math data, code data, etc.
Each sample contains one… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Domain_Translation.en-vi-translation
To join all training set files together
run python join_dataset.py file, final result will be join_dataset.json file
cantonese-mandarin-translations
Dataset Card for cantonese-mandarin-translations
Dataset Summary
This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese).
Supported Tasks and Leaderboards
N/A
Languages
Cantonese (yue)
Simplified Chinese (zh-CN)
Dataset Structure
JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.standard-malay-translation-instructionstranslation-checkpointsshp_translationsThis dataset contains translations of three splits (askscience, explainlikeimfive, legaladvice) of the Stanford Human Preference (SHP) dataset, used for training domain-invariant reward models.
The translation was conducted using the No Language Left Behind (NLLB) 3.3 B 200 model.
References:
Stanford Human Preference Dataset: https://huggingface.co/datasets/stanfordnlp/SHP
NLLB: https://huggingface.co/facebook/nllb-200-3.3B
Patent_Translation_datasetPatent Dataset ..
translation-quality
Multilingual Translation Quality Dataset
This dataset provides multilingual text chunks translated into English, accompanied by automated quality evaluations generated by multiple large language models.
Dataset Details
Source Data: agentlans/HuggingFaceFW-finetranslations-100-languages-sample
Target Language: English
Content: Multilingual chunks mapped to their English translations alongside automated judge scores.
Evaluation Methodology
The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/translation-quality.parallel-translation-training-pool
Parallel translation training pool
Sentences in eleven languages beside their translations, from five public parallel corpora read at
the pinned revisions named below and laid out twice. Ten languages are paired with English in both
directions, twenty directions in all. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 4975238 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/parallel-translation-training-pool.Code-Translationwellbeing-in-translation
Wellbeing in Translation
Raw outputs and translated materials for Does AI Wellbeing Survive Translation? We test whether the unchanged CAIS 1-7 self-report battery measures the same positive-minus-negative gap after translation.
Paper · Code · Source instrument
Headline result
Language sensitivity is specific to the model-battery pair.
Model
Gap spread across 7 languages
English rank
English stimulus / local battery
Local stimulus / English battery… See the full description on the dataset page: https://huggingface.co/datasets/ic-org/wellbeing-in-translation.parsinlu-machine-translation-en-fa-alpaca-style
ParsiNLU Machine Translation En-Fa in Alpaca Style
This dataset is an Alpaca-style and instruction-included version of the ParsiNLU original dataset.
HalluciGen-Translation
Task 2: HalluciGen - Tranlsation
This dataset contains the trial and test splits per language pair for the Translation scenario of the HalluciGen task, which is part of the 2024 ELOQUENT lab.
NOTE: A gold-labeled version of the dataset will be released in a new repository.
Dataset schema
id: unique identifier of the example
langpair: the source and target language pair of the example
source: original model input for translation
hyp1: first alternative translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/HalluciGen-Translation.ESFT-translationwikidata-entity-translationsEnglish-Hindi_Translation
📘 README.md
👉 Copy everything below into your repository README.md
English–Hindi Massive Synthetic Translation Dataset
🧠 Overview
This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines.
The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with:
100+ subjects
100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.sharegpt_deepl_ko_translationhttps://github.com/jwj7140/Gugugo
sharegpt_deepl_ko를 한-영 번역데이터로 변환한 데이터입니다.
translation_data_sharegpt.json: 최대 약 1300자 분량의 번역 데이터 모음
translation_data_sharegpt_long.json: 1300자~7000자 분량의 번역 데이터 모음
translation_data_sharegpt_long_newlineClean.json: translation_data_sharegpt_long.json에서 개행이 번역되지 않은 항목를 제거한 데이터
translation_data_sharegpt_long_newlineClean.json: translation_data_sharegpt.json에서 개행이 번역되지 않은 항목를 제거한 데이터
sharegpt_deepl_ko에서 몇 가지의 데이터 전처리를 진행했습니다.
optimal-reference-translationsThis is the dataset for two papers: Quality and Quantity of Machine Translation References for Automated Metrics [paper] - effect of reference quality and quantity on automatic metric performance, and Evaluating Optimal Reference Translations [paper] - creation of the data and human aspects of annotation and translation.
Please see the original repository for more information and the raw data or contact the authors with any questions.
Please make sure that you have the latest datasets… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/optimal-reference-translations.databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0.
Last Update : 2023-04-18
databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data
en-vi-translation-testDolci-Think-SFT-7B-translationsadaption-piguard-translation-handoff
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-piguard_translation_handoff
This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.welsh-translation-instructionThis is a set of Alpaca formatted Welsh-English translation instructions, obtained from the Welsh Government website.
translation-contrastive-tripletscherokee-english-translation
Cherokee–English Parallel Corpus (Archivist Project)
A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine
translation, assembled from public sources, deduplicated, benchmark-decontaminated,
and conflict-cleaned. Built to train and evaluate English→Cherokee translation
models for one of the most endangered languages in North America.
Files
File
Rows
Purpose
train_en2chr_v2.jsonl
138,307
Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.
