datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinh-phap-hoa-ke-trom-huongNormalized using https://github.com/oysterlanguage/emiliapipex
@article{emilia,
title={Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation},
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu, Zhizheng},
journal={arXiv},
volume={abs/2407.05361}… See the full description on the dataset page: https://huggingface.co/datasets/hr16/kinh-phap-hoa-ke-trom-huong.LibriSEA
LibriSEA 🎙️
LibriSEA is a multilingual speech dataset sourced from South East Asian audiobook recordings, designed for Automatic Speech Recognition (ASR), Text To Speech (TTS). Inspired by the LibriSpeech project, LibriSEA brings high-quality, read-speech data to underrepresented languages of Southeast Asia.
Dataset Summary
LibriSEA contains segmented audio clips extracted from publicly available Bible audiobooks across Southeast Asian languages. Each segment is… See the full description on the dataset page: https://huggingface.co/datasets/hr16/LibriSEA.
