datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
somali-combined-asr-stt-dataset
Somali Combined ASR/STT Dataset
A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining
synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript
text and split into train/validation/test.
Dataset Summary
Language: Somali (so)
Task: Automatic Speech Recognition / Speech-to-Text
Audio format: WAV, 16 kHz mono
Total examples: 8,226 (after deduplication)
Total audio: ~6 hours
Split
Examples
Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.fleurs-somali
FLEURS Somali
FLEURS Somali is a processed Somali speech dataset derived from the Somali portion of the Google FLEURS corpus. The dataset is designed for Automatic Speech Recognition (ASR), Speech-to-Text (STT), and speech technology research.
Audio samples have been enhanced through noise reduction and normalization while preserving the original speech content and transcriptions.
Features
Feature
Type
audio
Audio
transcription
String… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/fleurs-somali.
