CoolFace
Datasetpublicgated

proxectonos/Nos_Transcrispeech-GL

Corpus description Manually transcribed and speech-to-text aligned Galician ASR corpus containing 50 hours of multi-domain speech. The corpus contains different types of audios: conferences, debates, speeches, and interviews. The corpus is divided into three partitions: train (80%), dev(10%) and test (10%). Each partition contains audio fragments in WAV format, aligned with their respective transcriptions. In the "/data" folder you can find the audio partitions, while the "/transcript"… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Nos_Transcrispeech-GL.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes53downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.