datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cebuano-speech
Cebuano (Bisaya) Spontaneous Speech — Silencio Philippines Pack
Spontaneous long-form Cebuano with human transcription and word-level forced alignment. Fifteen speakers, mean clip length over two minutes, 27,000+ timestamped tokens. Part of the Silencio Philippines Pack.
Hours
3.48
Clips
90
Speakers
15
Countries
2
Speaker origin regions
4
L1 speakers of the recorded language
11 of 15 (65 clips)
Audio
48 kHz stereo WAV
Mean clip length
139.2 s… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/cebuano-speech.Cebuano-Speech-Dataset
🎧 Cebuano Speech Dataset
The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.
