laki35/somali-stt-dataset-multi-speaker-v1
Dataset Structure The dataset contains the following columns: text: The Somali sentence (transcription). audio: The audio file sampled at 24,000 Hz. speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers. Metadata & Search Keywords Language: Somali (so) Speakers: 11 unique voices (balanced gender representation) Audio Quality: 24kHz, mono, clean audio Total Rows: 1,200 Total Duration: ~1.66 Hours (99.86 Minutes) Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.
294
Dataset Structure
The dataset contains the following columns:
text: The Somali sentence (transcription).audio: The audio file sampled at 24,000 Hz.speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.
Metadata & Search Keywords
- Language: Somali (so)
- Speakers: 11 unique voices (balanced gender representation)
- Audio Quality: 24kHz, mono, clean audio
- Total Rows: 1,200
- Total Duration: ~1.66 Hours (99.86 Minutes)
- Intended Use: Fine-tuning multilingual TTS models (e.g., coqui/XTTS-v2 not from scratch)
