CoolFace
Datasetpublicgated

proxectonos/Nos_Transcrispeech-GL

Corpus description Manually transcribed and speech-to-text aligned Galician ASR corpus containing 50 hours of multi-domain speech. The corpus contains different types of audios: conferences, debates, speeches, and interviews. The corpus is divided into three partitions: train (80%), dev(10%) and test (10%). Each partition contains audio fragments in WAV format, aligned with their respective transcriptions. In the "/data" folder you can find the audio partitions, while the "/transcript"… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Nos_Transcrispeech-GL.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes61downloads

proxectonos/Nos_Transcrispeech-GL · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.