kudo-research/mustc-en-es-text-only
Dataset Card for kudo-research/mustc-en-es-text-only Dataset Summary This dataset is a selection of text only (English-Spanish) from the MuST-C corpus. MuST-C is a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 14 languages (Arabic, Chinese, Czech, Dutch, French, German, Italian, Persian, Portuguese, Romanian, Russian, Spanish, Turkish and Vietnamese). For each target… See the full description on the dataset page: https://huggingface.co/datasets/kudo-research/mustc-en-es-text-only.
062
