datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.eurospeech-bg-single-speaker
EuroSpeech BG — single-speaker subset
Bulgarian parliamentary speech from disco-eth/EuroSpeech,
filtered down to clips containing exactly one speaker.
Why
EuroSpeech ships no speaker labels — its only identity-like field, video_id,
is a parliamentary session containing dozens of speakers. To build
LibriSpeechMix-style simulated mixtures for speaker-diarization training you
first need clean single-speaker source audio. This subset is that source.
How… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-single-speaker.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.knesset-committees-speakers
Knesset Committees Speakers
An index that attaches a verified Knesset member identity, and through it
demographics, to the committee audio in
ivrit-ai/knesset-committees.
No audio is included. Each row names a span (session, start, end) of that
dataset's audio.m4a; filename follows the VoxKnesset convention
{speaker_id}_{session}_{start_ms}_{end_ms}.wav so the same tooling applies.
speaker_id is the Knesset's official PersonID -- the same id space as the
Knesset Corpus and… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/knesset-committees-speakers.Clean_One_Speaker_STT_Split_EN_AR
Clean_One_Speaker_STT_Split_EN_AR
Mixed STT dataset — Saudi dialectal Arabic + Modern Standard Arabic + English —
with train/validation/test splits, built for fine-tuning multilingual ASR (e.g. Whisper)
without catastrophic forgetting of English.
Composition
Dialectal Arabic (~72%) — Sebssihakim/Clean_One_Speaker_SADA_Split
(cleaned single-speaker SADA; SDAIA / Saudi Broadcasting Authority, CC BY-NC-SA 4.0)
MSA (~8%) — FLEURS ar_eg, full official splits (CC-BY;… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_STT_Split_EN_AR.synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.synthetic-speaker-diarization-dataset-fa-large-3000audio-3-speaker-dataset-v2
Three-Speaker Audio Dataset with Timbral Speaker Embeddings
A teaching dataset maintained by AI-Academy. It pairs single-speaker English
utterances with precomputed timbral speaker embeddings, and is intended for
coursework and exercises rather than for benchmarking or production systems.
The dataset deliberately contains one injected inconsistency; locating it is one of
the intended exercises (see The injected anomaly).
Overview
Property
Value
Examples… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/audio-3-speaker-dataset-v2.romanian-tts-single-speaker
Romanian TTS Single Speaker
A single-speaker Romanian speech dataset for TTS model training.
Dataset Description
Segments
24,379
Duration
34.3 hours
Speaker
Sanda (female)
Language
Romanian (ro)
Audio
WAV, 16-bit, mono, 24 kHz
Subsets
Subset
Segments
Description
standard
24,203
Standard Romanian sentences
loanword
176
Sentences containing foreign loanwords
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE
AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE
Overview
This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps.
The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.Clean_One_Speaker_SADA
Clean_One_Speaker_SADA
A cleaned, single-speaker subset of the SADA (Saudi Audio Dataset for Arabic) corpus,
derived from MahmoudIbrahim/100Hours-SADA22.
Processing
Follows the cleaning procedure from the Kaggle notebook
Segmented Audio Data for Arabic Dialects (SADA):
Removed rows with Unknown speaker age or gender.
Removed rows whose dialect is More than 1 speaker, Unknown, or Notapplicable
(every remaining segment has exactly one identified speaker).… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA.idoma-tts-speaker1📖 Abstract
Most speech technology exists for high-resource languages, leaving low-resource languages like Idoma digitally underrepresented. This project bridges that gap by developing a fully functional Bidirectional English-Idoma Speech System.
The system utilizes an Applied Research Design to solve data scarcity and tonal complexity issues inherent in the Idoma language. It integrates three fine-tuned transformer models to enable natural conversation flow:
ASR (Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/mrheartng/idoma-tts-speaker1.idoma-tts-multiple-speakers📖 Abstract
Most speech technology exists for high-resource languages, leaving low-resource languages like Idoma digitally underrepresented. This project bridges that gap by developing a fully functional Bidirectional English-Idoma Speech System.
The system utilizes an Applied Research Design to solve data scarcity and tonal complexity issues inherent in the Idoma language. It integrates three fine-tuned transformer models to enable natural conversation flow:
ASR (Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/mrheartng/idoma-tts-multiple-speakers.sada2022-speaker1-segments
SADA2022 Speaker1 Segments
This dataset contains cropped audio segments for the target speaker.
Format
Each shard follows AudioFolder-style layout:
shard_00001/
train/
metadata.csv
00000001.wav
00000002.wav
metadata.csv contains:
file_name
text
speakeroverlap_multiseg
MultiSeg Dataset for ASR Hallucinations
Description
MultiSeg is a perturbed and altered version of the TEDLIUM3 dataset, specifically created for evaluating the robustness of Automatic Speech Recognition (ASR) systems. This dataset is derived from the 'speakeroverlap' subset, which consists of held-back training data from TEDLIUM3.
Purpose
The primary purpose of the MultiSeg dataset is to:
Elicit hallucinations from ASR systems
Evaluate ASR performance under… See the full description on the dataset page: https://huggingface.co/datasets/zbrunner/speakeroverlap_multiseg.Clean_One_Speaker_SADA_Split
Clean_One_Speaker_SADA_Split
Train/validation/test version of
Sebssihakim/Clean_One_Speaker_SADA,
a cleaned single-speaker subset of the SADA corpus (SDAIA / Saudi Broadcasting Authority).
Splits
80/10/10, stratified by speaker_dialect, seed 42:
train ~30.8k, validation ~3.9k, test ~3.9k rows.
The Maghrebi dialect (7 rows) was removed — too few samples to stratify.
Splits are segment-level: the same source show/speaker may appear in more than one split.
Rare… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA_Split.sada2022-speaker1-segments
SADA2022 Speaker1 Segments
This dataset contains cropped audio segments for the target speaker.
Format
Each uploaded shard follows AudioFolder-style layout:
shard_xxxxx/
train/
00000001.wav
00000002.wav
metadata.csv
The metadata file contains:
file_name
text
