CMKL/Porjai-Thai-voice-dataset-central
Porjai-Thai-voice-dataset-central This corpus contains a officially split of 700 hours for Central Thai, and 40 hours for the three dialect each. The corpus is designed such that there are some parallel sentences between the dialects, making it suitable for Speech and Machine translation research. Our demo ASR model can be found at https://www.cmkl.ac.th/research/porjai. The Thai Central data was collected using Wang Data Market. Since parts of this corpus are in the ML-SUPERB… See the full description on the dataset page: https://huggingface.co/datasets/CMKL/Porjai-Thai-voice-dataset-central.
This repository belongs to CMKL on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
