CoolFace
Datasetpublic

suleiman2003/W_hausa_v3

Cleaned Hausa Speech Dataset v3 A cleaned and processed Hausa speech dataset built from multiple open-source Hugging Face datasets. Dataset Description This dataset contains cleaned, normalized, and deduplicated Hausa speech audio with aligned transcriptions. All audio is: Sample rate: 16,000 Hz (mono) Format: FLAC (lossless, embedded in Parquet) Duration range: 1–30 seconds per clip Loudness normalized: -20 dBFS RMS VAD trimmed: Non-speech segments removed with… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v3.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes329downloads

suleiman2003/W_hausa_v3 · main · files are served by the source, never re-hosted here