CoolFace
Datasetpublicgated

proxectonos/Nos_Transcrispeech-GL

Corpus description Manually transcribed and speech-to-text aligned Galician ASR corpus containing 50 hours of multi-domain speech. The corpus contains different types of audios: conferences, debates, speeches, and interviews. The corpus is divided into three partitions: train (80%), dev(10%) and test (10%). Each partition contains audio fragments in WAV format, aligned with their respective transcriptions. In the "/data" folder you can find the audio partitions, while the "/transcript"… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Nos_Transcrispeech-GL.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes61downloads
.gitattributesDownload Raw Back to root

This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.