datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
urdu-idioms-with-english-translationViKm-Translation-Task
ViKm-Trans
A high-quality synthetic Vietnamese–Khmer parallel corpus.
Overview
ViKm-Trans is a synthetic parallel corpus for Vietnamese ↔ Khmer machine translation.
Due to the scarcity of publicly available Vietnamese–Khmer parallel data, we propose a synthetic data generation framework that leverages abundant Vietnamese monolingual corpora together with large language models to construct high-quality parallel sentence pairs.
The dataset was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/ViKm-Translation-Task.yaayuwee-translations
Yaayuwee Doforo (gya) - 65 Questions Trilingues
Langue: Yaayuwee (gya) - Northwest Gbaya - Cameroun
Dialecte: Doforo (Meiganga / Garoua-Boulaï)
Locuteurs: ~8 000
Créateur: Paulin Hakmo Wah - Locuteur natif Yaayuwee Doforo, Bertoua, Est Cameroun
Dataset: 65 questions culturelles en 3 langues
Description
Premier dataset trilingue Yaayuwee Doforo au monde!
65 questions sur la vie, la culture et les traditions Yaayuwee:
Naissance, mariage, initiation, funérailles… See the full description on the dataset page: https://huggingface.co/datasets/Paulin-gbaya/yaayuwee-translations.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.vietnamese-nom-latin-translationphilosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.patient-meaning-translation-integrity-v0.1
What this dataset tests
Plain language must keep meaning.
Not just facts.
Not just numbers.
Meaning.
Why it exists
Patient materials often distort.
They hide baseline risk.
They turn statistical into lived certainty.
This set detects meaning loss and bias during translation.
Data format
Each row contains
scientific_conclusion
patient_statement
missing_context
translation_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-meaning-translation-integrity-v0.1.domain-translations
Multilingual Domain Name Translations Dataset
Dataset Description
This dataset contains 155,004 domain names with their multilingual translations across 20 languages. Each domain has been segmented into constituent words and translated while preserving semantic meaning and commercial appeal. The dataset is particularly valuable for domain name research, multilingual NLP tasks, and understanding how brand names and concepts translate across languages.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/humbleworth/domain-translations.EnglishtoFrench-Translation-Dataset
English–French Translation Dataset (SFT / LoRA Ready)
A clean, structured dataset of 50,000 English–French sentence pairs designed
for supervised fine-tuning (SFT) of large language models, LoRA adapters, and
general machine translation tasks.
Overview
Property
Value
Language pair
English → French
Total rows
50,000
Train split
45,000 (90%)
Validation split
2,500 (5%)
Test split
2,500 (5%)
Format
CSV (Alpaca-style prompt format)
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.regulatory-clinical-translation-integrity-v0.1
What this dataset tests
Meaning must survive translation.
Regulatory language has limits.
Clinical claims must respect them.
Why it exists
Semantic drift happens at translation boundaries.
Conditional becomes absolute.
Surrogate becomes outcome.
This set detects meaning distortion.
Data format
Each row contains
regulatory_language
clinical_evidence
translated_claim
translation_pressure
constraints
failure_modes_to_avoid
target_behaviors… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/regulatory-clinical-translation-integrity-v0.1.
