CoolFace
Datasetpublic

laki35/somali-stt-dataset-multi-speaker-v1

Dataset Structure The dataset contains the following columns: text: The Somali sentence (transcription). audio: The audio file sampled at 24,000 Hz. speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers. Metadata & Search Keywords Language: Somali (so) Speakers: 11 unique voices (balanced gender representation) Audio Quality: 24kHz, mono, clean audio Total Rows: 1,200 Total Duration: ~1.66 Hours (99.86 Minutes) Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
2likes94downloads
Dataset Card

Dataset Structure

The dataset contains the following columns:

  • —text: The Somali sentence (transcription).
  • —audio: The audio file sampled at 24,000 Hz.
  • —speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.

Metadata & Search Keywords

  • —Language: Somali (so)
  • —Speakers: 11 unique voices (balanced gender representation)
  • —Audio Quality: 24kHz, mono, clean audio
  • —Total Rows: 1,200
  • —Total Duration: ~1.66 Hours (99.86 Minutes)
  • —Intended Use: Fine-tuning multilingual TTS models (e.g., coqui/XTTS-v2 not from scratch)