datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.cross-species-translational-alignment
Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21
Goal: build a training substrate for detecting subtle / pre-histopathological
toxicity signatures in animal transcriptome data, with mechanism-of-toxicity
labels attached. This directory contains the compound-level linkage layer:
every compound that has rat in-vivo perturbation data cross-referenced to Tox21
mechanism assays via standardized chemical identifiers.
Background — the hackathon
Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tatoeba-english-translations
Tatoeba English Translation Dataset
Dataset Summary
This dataset is derived from the Tatoeba database, focusing on English sentences and their translations. It includes assessments of English sentences using text quality, sentiment, and readability models. The dataset is designed for tasks related to multilingual text quality, readability, and sentiment analysis.
Supported Tasks and Leaderboards
Quality Assessment
Readability Prediction
Sentiment Analysis… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/tatoeba-english-translations.Amazigh-Quran-Translation-Jouhadi
Dataset Card: Tamazight (Tifinagh) Quran Translation - Lahoucine Jouhadi
This dataset provides a digitized, partial translation of the meanings of the Holy Quran into Amazigh (Tachelhit) using the Neo-Tifinagh script. The content is based on the full translation work of Lahoucine Jouhadi (Lhocine Jouhadi Baamrani) based on Warsh recitation used in Morocco.
Original Sources & References
Author's Website - Down currently: Jouhadi Lahoussine Publications… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Quran-Translation-Jouhadi.cross_species_leaf_absolute_translationcross_species_leaf_on_off_translationEnglishtoFrench-Translation-Dataset
English–French Translation Dataset (SFT / LoRA Ready)
A clean, structured dataset of 50,000 English–French sentence pairs designed
for supervised fine-tuning (SFT) of large language models, LoRA adapters, and
general machine translation tasks.
Overview
Property
Value
Language pair
English → French
Total rows
50,000
Train split
45,000 (90%)
Validation split
2,500 (5%)
Test split
2,500 (5%)
Format
CSV (Alpaca-style prompt format)
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.Calibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
ASCAT-Arabic-Scientific-Translation
ASCAT: Arabic Scientific Corpus for Advanced Translation
ASCAT (Arabic Scientific Corpus for Advanced Translation) is a high-quality English–Arabic parallel corpus of full scientific abstracts designed for rigorous evaluation and training of domain-specific machine translation (MT) systems.
Unlike existing Arabic–English corpora that rely on short sentences or narrow domains, ASCAT targets long-form scientific abstracts validated through a multi-engine translation and expert… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/ASCAT-Arabic-Scientific-Translation.english-darija-arabizi-translationcrowdsourced-text-to-sign-language-rule-based-translation-corpus
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/sltAI/crowdsourced-text-to-sign-language-rule-based-translation-corpus.meps_speeches_with_translation.csvJeju-Standard-TranslationIrish_Translations
Irish_Focloir
This dataset contains English phrases along with their Irish translations.
Dataset Structure
The dataset contains the following fields:
english: The English phrase.
irish: The Irish translation of the phrase.
Example
Here is an example of the data structure:
english,irish
"Hello","Dia dhuit"
"Goodbye","Slán"
"Thank you","Go raibh maith agat"
wat2025-translation-collectionreferenceless_machine_translation_evaluationBengali is a low resource language in natural language processing (NLP), with dialects like Sylheti, Chittagong, and Barisal
being even more underrepresented. To address this, ONUBAD introduced a parallel corpus translating these dialects into
Standard Bangla and English using expert translators, providing 1,540 words, 130 clauses, and 980 sentences per dialect.
We focused on the Sylheti-English pair and adapted the dataset for LLM-based machine translation (MT) evaluation.
We extracted the… See the full description on the dataset page: https://huggingface.co/datasets/bokatiq/referenceless_machine_translation_evaluation.legal-scientific-evidence-translation-coherence-v0.1What this dataset is
You receive
scientific finding
court translation
uncertainty bounds
method limits
overstatement signals
You decide
Does the translation preserve the limits of the science
Answer
coherent
or
incoherent
Why this matters
When translation drifts
juries misread certainty
appeals rise
convictions or verdicts destabilise
translation_2wat24_text_to_text_translationCalibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
