datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr-evaluationsmtedx-v-eval
mTEDx-V Eval: Long-Form Multilingual X->En Speech Translation with Recoverable Video Evidence
Talk-level (long-form) evaluation manifests for visual-context-aware simultaneous
speech translation, built from the Multilingual TEDx (mTEDx)
corpus. Each record is one full TEDx talk whose talk_id is the real YouTube video ID,
so the original talk video can be obtained from its official source and frames can be
aligned to the sentence-level segment timestamps below (timestamps are on… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/mtedx-v-eval.eval
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: jonathan harrison
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/eval.
