CoolFace
Datasetpublic

maanka2/fleurs-somali

FLEURS Somali FLEURS Somali is a processed Somali speech dataset derived from the Somali portion of the Google FLEURS corpus. The dataset is designed for Automatic Speech Recognition (ASR), Speech-to-Text (STT), and speech technology research. Audio samples have been enhanced through noise reduction and normalization while preserving the original speech content and transcriptions. Features Feature Type audio Audio transcription String… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/fleurs-somali.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes38downloads
Dataset Card

FLEURS Somali

FLEURS Somali is a processed Somali speech dataset derived from the Somali portion of the Google FLEURS corpus. The dataset is designed for Automatic Speech Recognition (ASR), Speech-to-Text (STT), and speech technology research.

Audio samples have been enhanced through noise reduction and normalization while preserving the original speech content and transcriptions.

Dataset Description

  • —Creator: maanka2
  • —Source Dataset: Google FLEURS (Somali)
  • —Language: Somali (so)
  • —Sampling Rate: 16 kHz
  • —Speaker Type: Multiple speakers
  • —Task: Automatic Speech Recognition (ASR)

Features

FeatureType
audioAudio
transcriptionString

Processing Pipeline

The dataset was processed using:

  • —Noise Reduction
  • —Audio Normalization
  • —16 kHz Audio Standardization

These steps improve audio consistency and reduce background noise for ASR training.

Intended Uses

  • —Automatic Speech Recognition
  • —Speech-to-Text Systems
  • —Acoustic Model Training
  • —Speech Foundation Models
  • —Somali Language Technology
  • —Low-Resource Language Research

Dataset Structure

The dataset contains Somali speech recordings paired with their corresponding transcriptions.

Limitations

  • —Derived from the Somali subset of Google FLEURS.
  • —Recording conditions may vary across speakers.
  • —Not intended for speaker identification tasks.

Citation

If you use this dataset in research or production systems, please cite both this dataset and the original Google FLEURS dataset.

Acknowledgements

This dataset is based on the Google FLEURS Somali corpus and has been further processed to support Somali Automatic Speech Recognition research and development.