CoolFace
Datasetpublic

bilguun/ted_talks_en_mn_split

TED & TEDx Parallel Corpus (English-Mongolian) The dataset is composed of two distinct subsets: TED Talks (split en): English-language talks sourced from the official TED platform, paired with high-quality, human-generated Mongolian subtitles. TEDxUlaanbaatar (split mn): Mongolian-language talks from local TEDx events in Ulaanbaatar, paired with the original Mongolian subtitles and machine-translated English subtitles. This version of the dataset features segmented audio and… See the full description on the dataset page: https://huggingface.co/datasets/bilguun/ted_talks_en_mn_split.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes10downloads

bilguun/ted_talks_en_mn_split · main · files are served by the source, never re-hosted here