datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shp_translationsThis dataset contains translations of three splits (askscience, explainlikeimfive, legaladvice) of the Stanford Human Preference (SHP) dataset, used for training domain-invariant reward models.
The translation was conducted using the No Language Left Behind (NLLB) 3.3 B 200 model.
References:
Stanford Human Preference Dataset: https://huggingface.co/datasets/stanfordnlp/SHP
NLLB: https://huggingface.co/facebook/nllb-200-3.3B
DBNL-public-qa-english-translationcauldron_translations_nu
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zakcination/cauldron_translations_nu.mig-english-myanmar-translation
👨💻 English <--> Myanmar Translation Dataset
English and Myanmar (Burmese) နှစ်ဘာသာကို စုဆောင်းပေးထားသော dataset ဖြစ်ပါတယ်။
Machine Translation လုပ်ငန်းစဉ်မှာ တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 13911 ရှိပါတယ်။
How to use
from datasets import load_dataset
dataset = load_dataset("Ko-Yin-Maung/Eng2Mm-Translation")
output
DatasetDict({
train: Dataset({
features: ['input', 'output'],
num_rows: 12645
})… See the full description on the dataset page: https://huggingface.co/datasets/Ko-Yin-Maung/mig-english-myanmar-translation.quran-indonesia-tafseer-translationEnglish-Egyptian-Translation-finance
Bilingual Egyptian Finance Dataset
Dataset Description
This dataset contains bilingual text pairs in English and Egyptian Arabic focused on finance and financial topics. Each entry provides parallel translations covering various aspects of Egyptian and regional economics, making it valuable for translation models, multilingual NLP research, and economic analysis applications.
Key Features
Languages: English ↔ Egyptian Arabic (العامية المصرية)
Domain: Economics… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/English-Egyptian-Translation-finance.Deepseek-V3-Distilled-Ancient-Chinese-Translation
Dataset Card
这是一个文言文/白话文互译的高质量数据集,翻译精准,理由充分。总共有70K条数据,通过Deepseek-V3蒸馏获得。
Dataset Card Authors
Shen ZhuoKang From ECNU
Dataset Card Contact
10235101553@stu.ecnu.edu.cn
Translation
🇰🇿 Domain-Specific Translation, Kazakh Context
Dataset Summary
Domain-Specific Translation, Kazakh Context is a targeted dataset designed to improve the machine translation capabilities of Large Language Models (LLMs) between Kazakh and English.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
1,000
Total Words (approx.)
28,387
Avg. Words per Sample
28
Word Count Distribution (Per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Translation.Translation-Sample-Dataset
TsukiOwO/Translation-Sample-Dataset
Data Source
This dataset is a portion of open-r1/OpenR1-Math-220k.
Purpose
Serves as sample data for translation text.
