CoolFace
Datasetpublic

bilguun/ted_talks_en_mn_split

TED & TEDx Parallel Corpus (English-Mongolian) The dataset is composed of two distinct subsets: TED Talks (split en): English-language talks sourced from the official TED platform, paired with high-quality, human-generated Mongolian subtitles. TEDxUlaanbaatar (split mn): Mongolian-language talks from local TEDx events in Ulaanbaatar, paired with the original Mongolian subtitles and machine-translated English subtitles. This version of the dataset features segmented audio and… See the full description on the dataset page: https://huggingface.co/datasets/bilguun/ted_talks_en_mn_split.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes10downloads
Dataset Card

TED & TEDx Parallel Corpus (English-Mongolian)

The dataset is composed of two distinct subsets:

  • —TED Talks (split en): English-language talks sourced from the official TED platform, paired with high-quality, human-generated Mongolian subtitles.
  • —TEDxUlaanbaatar (split mn): Mongolian-language talks from local TEDx events in Ulaanbaatar, paired with the original Mongolian subtitles and machine-translated English subtitles.

This version of the dataset features segmented audio and text, with each segment having a maximum duration of 30 seconds. For the complete, unsegmented version, please refer to bilguun/ted_talks_en_mn.

Known Limitations

Machine Translation Quality: The primary limitation is the quality of the English translations in the TEDxUlaanbaatar split. As these are machine-generated, they may contain inaccuracies, grammatical errors, or mistranslations of nuanced or idiomatic expressions.

Subtitle Alignment and Errors: As the data is derived from subtitles, some entries may contain minor errors. This can include missing words or phrases, or slight mismatches between parallel sentences due to subtitle timing and segmentation. Users should consider a preprocessing step to handle potential misalignments.