datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
translations-raw
natgillin/translations-raw
Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines.
31,663 parquet files (1566.8 GB)
49 language pairs under data/<src-tgt>/
Schema: 5 columns — see below
Read-only for downstream pipelines. Do not delete or modify.
Schema
Each parquet has 5 columns:
column
type
description
source
string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.mala-bilingual-translation-corpus
MaLA Corpus: Massive Language Adaptation Corpus
This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources.
As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages).
Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.Weblate-Translations
Dataset Card for Weblate Translations
A dataset containing strings from projects hosted on Weblate and their translations into other languages.
Please consider donating or contributing to Weblate if you find this dataset useful.
To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this:
from datasets import load_dataset
dataset = load_dataset("ayymen/Weblate-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Weblate-Translations.NLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.translationOpenSubtitles_Translations_DatasetPontoon-Translations
Dataset Card for Pontoon Translations
This is a dataset containing strings from various Mozilla projects on Mozilla's Pontoon localization platform and their translations into more than 200 languages.
Source strings are in English.
To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this:
from datasets import load_dataset
dataset = load_dataset("ayymen/Pontoon-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Pontoon-Translations.Weblate-Translations
Dataset Card for Weblate Translations
A dataset containing strings from projects hosted on Weblate and their translations into other languages.
Please consider donating or contributing to Weblate if you find this dataset useful.
Dataset Details
Dataset Description
Curated by: Mohamed Aymane Farhi
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): Check the README YAML metadata… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Weblate-Translations.ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.Bambara-Speech-Translation-Data
AfVoices-Translated (Bambara-English)
This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks.
Methodology
We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository.
Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.Translations_Hungarian_public_websites
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982
Description
A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian
Citation
Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982
speech-translation-and-summarization
English-Centric Multilingual Audio Dataset
This dataset contains generated article and summary audio for English-centric multilingual directions.
Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits.
Included directions
amharic_english / english_amharic
arabic_english / english_arabic
bengali_english / english_bengali
chinese_simplified_english / english_chinese_simplified
english_english
french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.norsumm-nob-nno-translation
Nynorsk-Bokmål translation pairs
A multi-sentence parallel corpus of manual Nynorsk-Bokmål translations. These translations were extracted from the SamiaT/NorSumm dataset. You can read more about how the original dataset was created (including details about the manual translation process) in Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles by Samia Touileb et al..
Contact
David Samuel (davisamu@ifi.uio.no)… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norsumm-nob-nno-translation.instruction_translationsTranslation of Instruction datasetancient-modern_greek_translations
Dataset Card for Ancient-Modern Greek translations
The Ancient-Modern Greek translations dataset includes 100 sentences of Ancient Greek texts manually translated into Modern Greek. Original texts and translations have been extracted from the web sources cited below.
Δημοσθένους, Ὑπὲρ τῆς Ῥοδίων ἐλευθερίας (ell: Δημοσθένους, Υπέρ της Ελευθερίας των Ροδίων; eng: Demosthenes, On the Liberty of the Rhodians). Μτφρ. Β.Η. Τσακατίκας. χ.χ. Λόγοι του Δημοσθένη. Γ' Ολυνθιακός, Υπέρ της… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/ancient-modern_greek_translations.swe_smith_back_translationBack translate the swe-smith data to get the problem statment following the R2E sylte prompt. More details in https://github.com/SWE-bench/SWE-smith/issues/127
AraMix-Translation-Scores
AraMix-Translation-Scores
AdaMLLab/AraMix (minhash_deduped
subset, 178,883,241 rows) with a machine-translation-detection score added to every
document. All original columns are preserved.
Columns
column
type
description
id
string
unchanged from AraMix
source
string
unchanged from AraMix
text
string
unchanged from AraMix
mmbert_quality_score
float64
AraMix's original mmbert_score, renamed
mmbert_translated_score
float64
new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.Europarl-Translation-Instruct
Dataset Card for Europarl-Translation-Instruct
Waifu to catch your attention.
Dataset Details
Dataset Description
europarl-translation-instruct is a translation instruct dataset built from europarl data.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): English instruct (but various languages in)
License: cc-by-sa-4.0
Dataset Sources
Source Data: https://www.statmt.org/europarl/ (Transcript source)
Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.enwiki-translations英語WikipediaをLLMを用いて英日翻訳したデータセットです。
collectionサブセットは翻訳元となった英語Wikipediaの文章、datasetサブセットは英日翻訳されたテキストペアと翻訳元にした事例のcollectionサブセットにおけるidを収載したものです。
なお、出力の利用に際しては、翻訳に使用した各モデルの出力に関するライセンス規約に従ってください。
Phi3.5 MoEを利用して作成されたデータについては、Wikipediaのライセンスにしたがい、CC-BY-SA 4.0で利用可能であるものとします。
translation-analysis
NuBerea Translation Verse Texts
Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts.
Attribution
Upstream Data Sources
Source
License
Clementine Vulgate, NOCR
Public Domain
Luther Bible 1545, NOCR
Public Domain
Matthew's Bible 1537, Textus Receptus Bibles
Public Domain
NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.code-code-translation-java-csharp
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru
CodeXGLUE -- Code2Code Translation
Task Definition
Code translation aims to migrate legacy software from one programming language in a platform toanother.
In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.BenchMAX_General_Translation
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain.
We collect parallel test data from Flore-200, TED-talk, and WMT24.
Usage
Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.translation_sourcelior-translationundl_ru2en_translation
Dataset Card for "undl_ru2en_translation"
More Information needed
african_languages_translationrelaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.en-uk-translation-conversations
