datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MakeSense-Emilia-DatasetThis is a dataset of simultaneous interpretation policy / trajectory, includeing EN, ZH, JA, KO translation.
Each language has 2,000 records of asr data and 6,000 records of simultaneous interpretation trajectory data.
Source: amphion/Emilia-Dataset
Please refer to github to see the usage
Luo-Synthetic-ASR-DatasetSynthetic Dholuo ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 775 hours.
