datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.DBNL-public-qa-english-translationwork-translation
LLM repeated back-translation trajectories
What happens to a text when it is translated back and forth repeatedly by an LLM?
This exploratory dataset starts from short French source texts and sends them through different pivot languages, one translation at a time. Each translation step is a new, stateless API call: the model receives only the fixed translation instruction, the target language, and the previous step's text.
The dataset is designed to explore whether repeated… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/work-translation.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.oasst1_shan_translationThis datasets is a translation version of OpenAssistant/oasst1 to Shan language, translated by facebook/nllb-200-3.3B.
The data quality has not been checked by a human yet, so that it might be of low quality.
dclm-sample-13k-en-et-translationDocument level translations from English to Estonian derived from a sample from the DLCM dataset as present in the dolmino mix translated with google/gemma-3-27b-it .
EnglishtoFrench-Translation-Dataset
English–French Translation Dataset (SFT / LoRA Ready)
A clean, structured dataset of 50,000 English–French sentence pairs designed
for supervised fine-tuning (SFT) of large language models, LoRA adapters, and
general machine translation tasks.
Overview
Property
Value
Language pair
English → French
Total rows
50,000
Train split
45,000 (90%)
Validation split
2,500 (5%)
Test split
2,500 (5%)
Format
CSV (Alpaca-style prompt format)
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.
