datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
ashdatasets
Dataset Card for "ashdatasets"
More Information needed
anv-data-ke-somali-fullanv-data-ke-somali-fullSomali34minesMohamed-diirowsomali-combined-asr-stt-dataset
Somali Combined ASR/STT Dataset
A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining
synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript
text and split into train/validation/test.
Dataset Summary
Language: Somali (so)
Task: Automatic Speech Recognition / Speech-to-Text
Audio format: WAV, 16 kHz mono
Total examples: 8,226 (after deduplication)
Total audio: ~6 hours
Split
Examples
Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.somali-stt-dataset-multi-speaker-v1
Dataset Structure
The dataset contains the following columns:
text: The Somali sentence (transcription).
audio: The audio file sampled at 24,000 Hz.
speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.
Metadata & Search Keywords
Language: Somali (so)
Speakers: 11 unique voices (balanced gender representation)
Audio Quality: 24kHz, mono, clean audio
Total Rows: 1,200
Total Duration: ~1.66 Hours (99.86 Minutes)
Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.anv-data-ke-somalisomali-speech-datasetHussein_3somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.fleurs-somali
FLEURS Somali
FLEURS Somali is a processed Somali speech dataset derived from the Somali portion of the Google FLEURS corpus. The dataset is designed for Automatic Speech Recognition (ASR), Speech-to-Text (STT), and speech technology research.
Audio samples have been enhanced through noise reduction and normalization while preserving the original speech content and transcriptions.
Features
Feature
Type
audio
Audio
transcription
String… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/fleurs-somali.SomaliDatasetSomali-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali-Call-Center-Audio-Dataset-Single-Channel.somali_cleaned_datasetsomali-tts-datasetsanv-data-ke-somali-testSomali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali_Call_Center_Audio_Dataset_Dual_Channel.somali-whisper-master-corpus-5ksomali-tts-datasetafri-voices-somali-speechsom_agriashremovednoisy
Dataset Card for "ashremovednoisy"
More Information needed
anv-data-ke-somali-fullsomali_speechsomali_tts
Dataset Card for "somali_tts"
More Information needed
somali_cleaned_datasetarabic-somali-vocabb
