datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
translation-quality
Multilingual Translation Quality Dataset
This dataset provides multilingual text chunks translated into English, accompanied by automated quality evaluations generated by multiple large language models.
Dataset Details
Source Data: agentlans/HuggingFaceFW-finetranslations-100-languages-sample
Target Language: English
Content: Multilingual chunks mapped to their English translations alongside automated judge scores.
Evaluation Methodology
The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/translation-quality.wellbeing-in-translation
Wellbeing in Translation
Raw outputs and translated materials for Does AI Wellbeing Survive Translation? We test whether the unchanged CAIS 1-7 self-report battery measures the same positive-minus-negative gap after translation.
Paper · Code · Source instrument
Headline result
Language sensitivity is specific to the model-battery pair.
Model
Gap spread across 7 languages
English rank
English stimulus / local battery
Local stimulus / English battery… See the full description on the dataset page: https://huggingface.co/datasets/ic-org/wellbeing-in-translation.adaption-piguard-translation-handoff
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-piguard_translation_handoff
This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.cherokee-english-translation
Cherokee–English Parallel Corpus (Archivist Project)
A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine
translation, assembled from public sources, deduplicated, benchmark-decontaminated,
and conflict-cleaned. Built to train and evaluate English→Cherokee translation
models for one of the most endangered languages in North America.
Files
File
Rows
Purpose
train_en2chr_v2.jsonl
138,307
Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.bhagavad-gita-verses-sanskrit-translations
Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset
A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes:
📜 Original Sanskrit slokas
🔡 IAST transliteration
🌐 Multiple English & Hindi translations
🧠 Traditional commentaries from many teachers
🔢 Structured metadata (chapter, verse, IDs, authors)
This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.English_Arabic_Translation_Pairs
English · العربية
English→Arabic Technical & Reasoning Translation Dataset
High-quality English → Modern Standard Arabic translation pairs focused on
native-English educational, scientific, and reasoning content. English source
text is drawn from real corpora (FineWeb-Edu and three NVIDIA reasoning datasets);
Arabic translations are produced by DeepSeek-v4-flash under a strict
translation-only prompt that preserves notation, numbers, formulas, code, and
citations.
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/English_Arabic_Translation_Pairs.zh_translation_benchmark
zh_translation_benchmark
zh_translation_benchmark is a 2,000-example synthetic benchmark for evaluating whether a translation or rewriting model can produce natural Taiwan Traditional Chinese (zh-TW) from English, Mainland Chinese, Hong Kong Traditional Chinese, Cantonese-style written Chinese, or code-mixed English/Chinese documents.
The dataset now exposes a single Hugging Face subset/config: full. It is not split into dev and test files. The full 2,000 rows are loaded… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/zh_translation_benchmark.en-translations
Multilingual Parallel Sentences with Semantic Similarity Scores and Quality Metrics
This dataset is a diverse collection of parallel sentences in English and various other languages, sourced from multiple high-quality datasets.
Each sentence pair includes a semantic similarity score calculated using the Language-agnostic BERT Sentence Embedding (LaBSE) model,
along with additional quality metrics.
Supported Tasks
This dataset supports:
Machine Translation… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-translations.lmsys_persian_translationdpo_irish_eng_translationsThis is a test for my DPO dataset for Irish ENglish trasnlslations, raw data origin : https://www.gaois.ie/en/corpora/parallel?Query=Apple&Language=en&SearchMode=exact&PerPage=50, used COMETXL refrernce free maodel Unbabel/wmt23-cometkiwi-da-xl
(which has been trained to asses Irish) to score accepted/rejected. Used GPT4 to generate translations to compare with human stranslations of Irish legislation (which has to have a Irisng/English copy by law)
ImageCaptions-7M-Translations-Arabic-subset-150000Irish_English_Translation
Irish English graded Translations
Data collected in order to fine-tune a Irish-English LLM, see blog post for more details.
See here for the next phase, preference dataset formated (for DPO) here
Data Sources:
translated_gaois_graded.jsonl
Parallel English-Irish corpus of legislation collected by Gaois. This corpus contains high-quality, human-translated paragraph pairs, making it a valuable resource.
translated_tatoeba_graded.jsonl
One draw-back is uses a lot of… See the full description on the dataset page: https://huggingface.co/datasets/c123ian/Irish_English_Translation.bombo_hil_eng_raw_translationsmaasai-translation-corpus
Maasai-English Translation Corpus
Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally grounded tooling.
Overview
Total pairs: 9,910
Splits: 8,434 train / 738 valid / 738 test
Directions: 4,955 en→mas and 4,955 mas→en
Quality tiers: 8,444 gold and 1,466 silver
Main sources: 8,444 Bible-derived pairs, 680 cultural manual pairs, 70 knowledge-driven cultural pairs, 132 public-domain Hollis proverb pairs, 504 public-domain… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/maasai-translation-corpus.chem-machine-translation-samplesultrafeedback_translation_refined
