datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laos-voice-dataset-v2lao_stt_training_data
Lao Speech-to-Text Training Data
ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ.
ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure)
Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv
Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv
ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV:
audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.lao-speech-datasetlao-asr-thesis-datasetseniortalk
SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
Introduction
SeniorTalk is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on elderly aged 75 to 85. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), speaker dirazation (SD), speech editing and other… See the full description on the dataset page: https://huggingface.co/datasets/laolaiduowangshi/seniortalk.laos-speech-datasetlao-data-speechlao_asr_mergedlaos-speech-datasetlaoruiyalao-speech-dataset
