datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maleo-short-1.5H
Dataset Card for Maleo Short 1.5H
Dataset Description
Dataset Summary
Maleo Short 1.5H is a manually curated, rigorously annotated speaker diarization dataset designed to benchmark State-of-the-Art (SOTA) models against complex, "in-the-wild" media domains. While modern diarization pipelines excel in controlled acoustic environments (like telephony or reading corpora), they heavily struggle with the overlapping speech, sound effects, and rapid speaker shifts… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/maleo-short-1.5H.arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.emirates-dialect-speech-male
🌍 Emirates Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
📌 Source Data & Provenance
Source Repository: https://github.com/MahaAlBlooki/alsanaa-emirati-dataset
Domain & Content: Spoken Emirati dialectal Arabic speech recordings.
Dialect Focus: Emirates / Gulf Dialectal Arabic.
Standardized Format: 22,050 Hz… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/emirates-dialect-speech-male.openslr42-khmer-male
OpenSLR SLR42 Khmer Male Speech
This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset.
Dataset Description
This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions.
Each example contains:
audio: Khmer speech recording
text: Khmer transcription
Dataset Structure
Column
Type
Description
audio
Audio
Khmer speech recording
text
String
Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.ka-geo-voice-male-v1
Dataset Card for Georgian Male Voice Dataset v1
Intended Use
Primary Use: Training and fine-tuning TTS models for Georgian language synthesis, including microsoft/speecht5_tts.
Secondary Use: Research in speech synthesis, voice conversion, or linguistic analysis.
SpeechT5 Compatibility
This dataset is specifically formatted to be compatible with microsoft/speecht5_tts fine-tuning. The dataset includes:
Audio: 22,050 Hz mono WAV files (matching SpeechT5… See the full description on the dataset page: https://huggingface.co/datasets/akalandia/ka-geo-voice-male-v1.
