kudo-research/mustc-en-es-text-only
Dataset Card for kudo-research/mustc-en-es-text-only Dataset Summary This dataset is a selection of text only (English-Spanish) from the MuST-C corpus. MuST-C is a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 14 languages (Arabic, Chinese, Czech, Dutch, French, German, Italian, Persian, Portuguese, Romanian, Russian, Spanish, Turkish and Vietnamese). For each target… See the full description on the dataset page: https://huggingface.co/datasets/kudo-research/mustc-en-es-text-only.
Fix language tags (#3)
Fix `license` metadata (#2)
Update README.md (#1)
Upload dataset_infos.json
Upload data/test-00000-of-00001.parquet with git-lfs
Upload data/dev-00000-of-00001.parquet with git-lfs
Upload data/train-00000-of-00001.parquet with git-lfs
Create README.md
initial commit
initial commit
