datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hyvoxpopuli
HyVoxPopuli
HyVoxPopuli is an open Armenian speech dataset (~6 hours, 16 kHz mono) with raw and normalized transcripts. It targets ASR and TTS research for Eastern Armenian (hy-AM).
Note: Despite the name, this release is not the official Facebook VoxPopuli parliament corpus. Audio is literary narration (two voice actors) segmented into short clips. Update citations and experiments accordingly.
Dataset summary
Rows
623
Train / Val / Test
498 / 62… See the full description on the dataset page: https://huggingface.co/datasets/Edmon02/hyvoxpopuli.RTPSpeechPortuguese Speech
This dataset aims to provide people with European Portuguese audio and textual data pairs, which can be used to fine-tune large language models.
These Portuguese recordings are from RTP (Rádio e Televisão de Portugal), which we have broken down and transcribed into short sentences.
Citation
If you use this dataset, please cite:
L. M. Hoi, Y. Sun and S. K. Im, "An Automatic Speech Segmentation Algorithm of Portuguese based on Spectrogram Windowing," 2022 IEEE World AI IoT… See the full description on the dataset page: https://huggingface.co/datasets/edmond5995/RTPSpeech.
