CoolFace
Datasetpublic

Blakus/Spanish_spain_dataset_100h

Total hours: 120h. Language: Spanish. Source: Librivox (https://librivox.org/search?primary_key=5&search_category=language&search_page=1&search_form=get_results) Number of speakers: 17 without counting the collaborative audio books. Collection: Cutted by the windows speech recognition using the source text as grammars, then validated with Deep Speech Spanish model. Type of speech: Clean speech. Collected by: Carlos Fonseca M @ https://github.com/carlfm01 License : Public Domain Quality: low… See the full description on the dataset page: https://huggingface.co/datasets/Blakus/Spanish_spain_dataset_100h.

sourceHugging Facepddlupdated 2y agoView on Hugging Face
1likes15downloads
Dataset Card

Total hours: 120h.

Language: Spanish.

Source: Librivox (https://librivox.org/search?primarykey=5&searchcategory=language&searchpage=1&searchform=get_results)

Number of speakers: 17 without counting the collaborative audio books.

Collection: Cutted by the windows speech recognition using the source text as grammars, then validated with Deep Speech Spanish model.

Type of speech: Clean speech.

Collected by: Carlos Fonseca M @ https://github.com/carlfm01

License : Public Domain

Quality: low (16 Khz)

This dataset consists of a WAVs folder with the audios, plus a .txt file with the path to the audio and the speaker's transcription.