proxectonos/Nos_Transcrispeech-GL
Corpus description Manually transcribed and speech-to-text aligned Galician ASR corpus containing 50 hours of multi-domain speech. The corpus contains different types of audios: conferences, debates, speeches, and interviews. The corpus is divided into three partitions: train (80%), dev(10%) and test (10%). Each partition contains audio fragments in WAV format, aligned with their respective transcriptions. In the "/data" folder you can find the audio partitions, while the "/transcript"… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Nos_Transcrispeech-GL.
053
No card is published for this repository, or it could not be fetched from Hugging Face right now.
