CoolFace
Datasetpublic

kudo-research/mustc-en-es-text-only

Dataset Card for kudo-research/mustc-en-es-text-only Dataset Summary This dataset is a selection of text only (English-Spanish) from the MuST-C corpus. MuST-C is a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 14 languages (Arabic, Chinese, Czech, Dutch, French, German, Italian, Persian, Portuguese, Romanian, Russian, Spanish, Turkish and Vietnamese). For each target… See the full description on the dataset page: https://huggingface.co/datasets/kudo-research/mustc-en-es-text-only.

sourceHugging Facecc-by-nc-nd-4.0updated 4y agoView on Hugging Face
0likes62downloads
10 commits on main
c03311f4y ago

Fix language tags (#3)

albertvillanova
92f8db44y ago

Fix `license` metadata (#2)

albertvillanova, julien-c
3ef44894y ago

Update README.md (#1)

dave-kudo, lbourdois
0a2c25a5y ago

Upload dataset_infos.json

dave-kudo
4e9e4135y ago

Upload data/test-00000-of-00001.parquet with git-lfs

dave-kudo
71d51705y ago

Upload data/dev-00000-of-00001.parquet with git-lfs

dave-kudo
cd16d825y ago

Upload data/train-00000-of-00001.parquet with git-lfs

dave-kudo
6c94d9c5y ago

Create README.md

dave-kudo
72cdd2c5y ago

initial commit

dave-kudo
a098b275y ago

initial commit

system