datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghana-female-twi-speech-asr-8word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
51139 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-8word-splits.ghana-female-twi-speech-asr-full-length
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Audio-text dataset with 76 pairs of Twi (Ghanaian language) speech data.
Structure
audio/ - WAV audio files ({len(pairs)} files)
text/ - Corresponding text transcripts ({len(pairs)} files)
dataset_manifest.json - Links audio to… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-full-length.ghana-female-twi-asr-16word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-asr-16word-splits.ghana-female-twi-8sec-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-8sec-splits.saudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.dahab-egyptian-female-tts
Dahab — Egyptian Arabic, single female speaker
134.7 hours across 59,505 clips of Egyptian (Cairene) Arabic from one
female speaker, at 24 kHz mono. 26,741 clips (44.9%) carry diacritized
transcripts. Built for TTS fine-tuning.
Segmented from a single YouTube cooking channel, so the register is
conversational instructional speech throughout.
Structure
The train split is stored in self-contained Parquet shards. Each row contains an audio object with embedded WAV… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/dahab-egyptian-female-tts.excited-english-female-vocal-dataset
Micro-Expression Vocal Dataset Matrix (100-Word Operational Toolkit)
Brand Engine: Marie DeVox
Maintainer Intake: mariedevox@voicevendor.store | https://payhip.com/MarieDeVox
Format Architecture: Uncompressed WAV (44.1kHz / 24-bit Linear PCM)
Metadata Schema: LJ Speech-Compliant Structural Mapping (metadata.csv)
IMPORTANT
The audio assets and metadata files in this directory are a free preview intended… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/excited-english-female-vocal-dataset.
